back

by rramadass·29d ago·view on hn ↗
People seem to be making all sorts of assumptions/conjectures about the project and confusing themselves without reading the paper explaining the project (linked to in the site itself);

From the introduction;

Classical Sanskrit recitation, parāyaṇa, is a chanted rather than a read register. A faithful synthesizer must hold long vowels, sustain a terminal visarga, articulate retroflex and aspirated consonants, render dense consonant conjuncts cleanly, and respect the metrical structure of the verse.

None of these is well served by general-purpose text-to-speech, and there is essentially no chant-domain training data available off the shelf. The problem is therefore doubly hard: it is low-resource, and the target prosody is a specialized melodic contour rather than ordinary read speech.

This report describes a system, Vāgdhenu, that solves the practical version of this problem well enough to ship two large deployments, and it documents the design decisions, the dead ends, and the one negative result that turned out to be the most useful finding. We do not claim a new model. We claim an honest account of what it takes to build a faithful Sanskrit chant pipeline on top of current open backbones, what works, and what is architecturally out of reach.

Our framing is that of an experience report. The evidence we offer is the comparative lineage across architecture families, a reproducible production system, two shipped artifacts at real scale, and a public release of code, weights, data, and a live demonstration. Formal listening studies are limited to expert evaluation, which we state plainly and treat as a limitation rather than a result.

1 comments
Perhaps you could edit the submission to add this context?
Sorry, am not the submitter so we will let it be.