A studio-grade voice corpus
for the Urdu-speaking world.
230 million Urdu speakers. Zero gold-standard voice data. We are building the dataset and the models that let modern AI finally speak the language the way it is actually spoken — in Karachi, in Lucknow, in Lahore, in the diaspora.
What studio-grade Urdu voice sounds like.
A preview of the voices, dialects and registers inside the corpus. Full samples arrive as recording wraps on each cohort.
A corpus, a model, a platform.
The Corpus
4,200 hours across 1,500 speakers. Spanning Karachi, Lahori, Lucknowi, Deccani, Peshawari and diaspora voices. Studio-grade for TTS, broadcast-grade for ASR, consented end-to-end.
The Models
TTS, voice cloning and ASR built on the corpus. Nastaliq-aware, code-switch tolerant, dialect-conditioned. Benchmark-first — every release ships numbers, not slogans.
The Platform
API access for product builders. Dataset licensing for AI labs. Research grants for linguists working on under-resourced Urdu phenomena.
Hear the first voices before anyone else.
Join the waitlist. We will send a short note when the first voice samples, dataset previews and model demos go live. No spam, unsubscribe any time.