Ancient texts. Living speech.
Arriving soon.
VĀṆĪ is an open corpus for Indian language research — classical Sanskrit and Tamil literature alongside living, code-switched speech data. We're preparing the first release, drawn from the Vedas to the Sangam poets to real conversations recorded with consent on Vayu.
No spam, one email at launch. Or write to us directly at corpus@vanic.org.
Real-world Indian language audio and transcripts, captured with consent via Vayu. English, Hindi, Tamil, Hinglish.
60+ classical Sanskrit and Tamil texts, annotated in a unified NLP-ready format — from the Rigveda to the Tirukkuṟaḷ.