How it's made
How Bulbul's Urdu audio is made
One narrator, two speeds, a review before anything ships, and no child's name ever sent to a voice service. Here's exactly how the Urdu you hear in Bulbul is produced.
Last updated 3 October 2026
Hear it
Three lines from the first lessons, as they sound in the app. Tap Slower to hear the same clip slowed down.
السلام علیکم
Assalam-o-Alaikum
Peace be upon you (hello)
MP3آپ کیسے ہیں؟
Aap kaise hain?
How are you?
MP3مجھے پانی چاہیے
Mujhe paani chahiye
I want water
MP3
These clips come from the app's own catalogue and play in your browser. Here, Slower slows the clip on your device; in the app, Slower plays a separately performed slow take once one is approved, and the normal take slowed until then.
Who is the voice in Bulbul?
One warm, clear narrator reads every line in Bulbul: the Urdu words and sentences, the English meanings, Bulbul's cheers, and the lines of family characters like Nani and Dada. The voice is called Reva. It's a voice from ElevenLabs' library, chosen for natural Pakistani Urdu pronunciation, and each line is generated from the written Urdu with ElevenLabs' text-to-speech (the Eleven v4 model).
One voice for every character was a deliberate choice. Children hear Ammi's line and Dada's line in the same voice, which keeps lessons calm and predictable, and it means there is never a synthetic child's voice in the app.
Is the audio recorded by a real person?
No. Bulbul's lines are synthesised from the written Urdu with ElevenLabs' AI voice technology, not recorded in a studio. We describe it as warm, clear Urdu narration, and we don't describe it as recordings of a person. If a word ever sounds off to you, email us: corrections go back through the same review as everything else.
Why a synthetic narrator?
- Consistency. Every word, in every lesson, in the same voice, at the same warmth, with the same pronunciation, including the family characters' lines.
- Every word, both speeds. The course has hundreds of lines. A new word or a corrected sentence can be produced, listened to and reviewed quickly, in a normal and a slower version, rather than waiting for a studio session.
- No children's voices. We never synthesise or record a child's voice. Bulbul's cheers, like “Shabash!” and “Bohat acha!”, are the same narrator.
What we don't claim. We don't claim the narration is a person's recording, that it has been checked by an outside reviewer yet, or that it's a model of perfect pronunciation. A pronunciation review is part of the release checklist, and we'll say so here once it has happened.
Is there a slower version of every word?
Yes. Every word and sentence in Bulbul has a Slower button next to Hear Urdu. Bulbul's audio contract asks for two files for every line, a normal take and a separately performed slow take, and the release check fails if either is missing.
Today the slower option is made by playing the normal take at a slower speed on the device, with the pitch kept. Separately performed slow takes are part of the release checklist, and the app uses one as soon as it exists.
What is reviewed before a line ships?
- The written Urdu. Each line is authored in Urdu script with its Roman spelling and English meaning. Changing the wording changes the line's identity, so an approved clip can't drift away from its text.
- The sound. A designated Urdu-speaking reviewer listens to each clip in context, normal and slow, for pronunciation, naturalness, family terms, polite forms and nasal sounds, then approves the exact file.
- The file. The approval record stores a fingerprint (a SHA-256 hash) of the approved text and of the approved audio file. Before a release is signed, the app's build pipeline checks every line against that record and refuses to build with a missing, changed or unapproved clip.
Where things stand on : the review process and the release check are built, and the approval record is still empty. The audio in Bulbul's development builds, and the clips on this website, were generated with the narrator above and listened to by us, but not yet formally approved. No line ships in a release until it is.
Does Bulbul ever send my child's name to a voice service?
Never. The narrator can't say every child's name, and the app never synthesises, bundles or fetches a name. In the line Mera naam … hai (My name is …), the narrator says “Mera naam”, leaves a pause where the name goes, and says “hai”. When it's your child's turn, they say the whole line with their own name, and the speech check accepts any name in that slot. Only the text on screen shows the name.
Where does the speech check run?
On the device. When your child speaks, Bulbul keeps the try in a temporary file on the iPhone or iPad and deletes it when the lesson ends or the app closes. Where iOS provides Apple's on-device Urdu speech recognition, Bulbul checks the try there: nothing is uploaded, and no outside service is involved. Apple's Urdu model is a one-time download of several hundred megabytes that starts by itself on Wi-Fi.
On devices without it, Bulbul listens and says “I heard you” without checking. Bulbul only ever checks for Urdu; it never uses another language's recogniser as a stand-in. The check means “Bulbul understood”, nothing more: it doesn't grade pronunciation, and it never counts toward progress.
Can I switch the speech check off?
Yes. In the Parent Area, behind the grown-up gate, there's a switch called “Bulbul checks spoken Urdu”for each learner. Off, Bulbul still listens hands-free and says “I heard you”, without checking the Urdu. The same screen tells you plainly what the check can do on your device, and whether Apple's Urdu model is downloaded.



