Awaz Media · The voice of Urdu, for the AI era

A studio-grade voice corpus
for the Urdu-speaking world.

230 million Urdu speakers. Zero gold-standard voice data. We are building the dataset and the models that let modern AI finally speak the language the way it is actually spoken — in Karachi, in Lucknow, in Lahore, in the diaspora.

01 · Hear it

What studio-grade Urdu voice sounds like.

A preview of the voices, dialects and registers inside the corpus. Full samples arrive as recording wraps on each cohort.

Preview coming
Lucknowi Classical ghazal
ہزاروں خواہشیں ایسی کہ ہر خواہش پہ دم نکلے
Hazaaron khwaahishen aisi, keh har khwaahish pe dam nikle
0:12
Preview coming
Karachi Conversational
آج شہر میں ٹریفک بہت زیادہ ہے، نکلنے سے پہلے پوچھ لینا
Aaj sheher mein traffic bohat zyada hai, nikalne se pehle pooch lena
0:08
Preview coming
Lahori Broadcast news
اقوام متحدہ کے سربراہ نے آج ایک اہم اعلان کیا
Aqwaam-e-mutahida ke sarbarah ne aaj ek ahem elaan kiya
0:10
02 · What we're building

A corpus, a model, a platform.

The Corpus

4,200 hours across 1,500 speakers. Spanning Karachi, Lahori, Lucknowi, Deccani, Peshawari and diaspora voices. Studio-grade for TTS, broadcast-grade for ASR, consented end-to-end.

The Models

TTS, voice cloning and ASR built on the corpus. Nastaliq-aware, code-switch tolerant, dialect-conditioned. Benchmark-first — every release ships numbers, not slogans.

The Platform

API access for product builders. Dataset licensing for AI labs. Research grants for linguists working on under-resourced Urdu phenomena.

03 · Early access

Hear the first voices before anyone else.

Join the waitlist. We will send a short note when the first voice samples, dataset previews and model demos go live. No spam, unsubscribe any time.