How Well Does Microsoft Speech-to-Text Understand Indian Speech?
Results from the Voice of India Benchmark — testing Microsoft Speech-to-Text across 8 of the benchmark's 15 languages.
Microsoft is scored on eight of fifteen languages.
Microsoft Speech-to-Text does not support these seven languages:
- Assamese
- Gujarati
- Kannada
- Maithili
- Odia
- Punjabi
- Telugu
Every figure in this report is confined to the eight it does cover.
| Microsoft is scored on these eight | Microsoft does not support these seven | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| System | Hindi | Bengali | Marathi | Tamil | Malayalam | Bhojpuri | Chhattisgarhi | Urdu | Assamese | Gujarati | Kannada | Maithili | Odia | Punjabi | Telugu | |
| Indic Transcribe Core | 3.5 | 4.3 | 5.7 | 9.0 | 11.5 | 13.3 | 13.6 | 5.0 | 8.2 | 9.2 | 7.4 | 11.3 | 8.7 | 8.3 | 11.4 | |
| Saaras v3 | 3.8 | 5.2 | 6.5 | 9.1 | 12.2 | 17.9 | 14.0 | 7.5 | 9.1 | 9.7 | 8.8 | 14.2 | 11.1 | 8.6 | 13.5 | |
| Gemini 3 Pro | 4.7 | 6.7 | 8.8 | 11.7 | 16.3 | 15.3 | 13.4 | 6.9 | 17.0 | 13.4 | 14.0 | 20.3 | 17.9 | 12.7 | 18.4 | |
| ElevenLabs Scribe v2 | 5.9 | 7.9 | 10.3 | 15.3 | 16.4 | 17.6 | 14.2 | 19.2 | 11.8 | 18.0 | 13.9 | – | 17.7 | 13.2 | 19.6 | |
| IndicConformer | 6.1 | 8.9 | 10.7 | 16.0 | 20.9 | 30.1 | 24.4 | 7.4 | 10.8 | 15.5 | 16.0 | 15.4 | 12.2 | 12.8 | 19.8 | |
| Gemini 3 Flash | 7.0 | 11.0 | 14.1 | 17.3 | 23.0 | 20.1 | 21.6 | 10.9 | 24.9 | 20.2 | 18.4 | 27.9 | 23.2 | 17.8 | 25.4 | |
| Microsoft Streaming STT | 7.9 | 13.6 | 26.9 | 22.1 | 33.4 | 31.8 | 25.6 | 17.1 | – | – | – | – | – | – | 17.2 | |
| Microsoft STT | 7.5 | 23.3 | 29.3 | 22.8 | 35.7 | 31.6 | 26.2 | 23.3 | – | – | – | – | – | – | – | |
| Gemma E4B | 9.1 | 19.6 | 24.4 | 37.9 | 44.4 | 27.0 | 24.0 | 14.2 | 45.0 | 27.3 | 31.0 | 36.7 | 44.4 | 23.4 | 41.6 | |
| GPT Realtime | 10.3 | 15.8 | 20.5 | 23.6 | 37.6 | 31.2 | 27.3 | 39.7 | – | 26.2 | 32.9 | – | – | – | 30.0 | |
| OmniASR LLM 7B v2 | 9.6 | 20.9 | 24.5 | 40.6 | 48.8 | 26.3 | 20.7 | 14.8 | 23.9 | 31.9 | 35.0 | 44.6 | 72.3 | 31.8 | 48.6 | |
| Amazon Transcribe | 4.8 | 6.5 | 7.7 | 12.3 | 20.6 | 27.5 | 22.4 | – | – | 13.8 | 10.9 | – | 13.8 | 12.3 | 13.9 | |
| Deepgram Nova 3 | 9.3 | 27.1 | 42.1 | 66.4 | – | 39.7 | 34.0 | – | – | – | 51.0 | – | – | – | 41.1 | |
| Ringg | 12.0 | – | – | – | – | – | – | – | – | – | – | – | – | – | – | |
Findings across eight languages
TEN KEY FINDINGSOur re-measurement, eight languages
| Language | Microsoft STT | Saaras v3 | Gemini 3 Pro | ElevenLabs Scribe v2 | IndicConformer | Amazon Transcribe | Gemini 3 Flash |
|---|---|---|---|---|---|---|---|
| Hindi | 7.5 | 3.8 | 4.7 | 5.9 | 6.1 | 4.8 | 7.0 |
| Bengali | 23.3 | 5.2 | 6.7 | 7.9 | 8.9 | 6.5 | 11.0 |
| Marathi | 29.3 | 6.5 | 8.8 | 10.3 | 10.7 | 7.7 | 14.1 |
| Tamil | 22.8 | 9.1 | 11.7 | 15.3 | 16.0 | 12.3 | 17.3 |
| Malayalam | 35.7 | 12.2 | 16.3 | 16.4 | 20.9 | 20.6 | 23.0 |
| Bhojpuri | 31.6 | 17.9 | 15.3 | 17.6 | 30.1 | 27.5 | 20.1 |
| Chhattisgarhi | 26.2 | 14.0 | 13.4 | 14.2 | 24.4 | 22.4 | 21.6 |
| Urdu | 23.3 | 7.5 | 6.9 | 19.2 | 7.4 | – | 10.9 |
When nobody said a time, it writes a clock anyway.
A clock-time opportunity is a clip whose reference contains no time word and no clock — a clip in which nobody said a time. There are 154,555 of them across the eight languages this system covers. It writes a colon-formatted clock on 3,290 of them. The six scored comparators, on the same clips, write one 0 times.
It is not a mis-hearing. A number word is present in the audio — “two”, “a”, “a minute” — and a post-processing stage decides that a number standing near another number must be a time, and formats it as one. The minutes it prints were never spoken and appear in no accepted spelling of the reference.
Microsoft writes 3,244. The next highest is 20 — 124× lower.
15 of the 23 systems never do it once. Microsoft's own streaming product is one of them: it fabricates a clock zero times on the same clips.
Ordinary words come back as digits.
Microsoft's model keeps returning digits in sentences where no number exists. These are not misread numbers — they are ordinary words, replaced by numerals. On this test it happens 6,028 times. Across the six comparators combined, it happens 4 times.
"No number exists" is meant literally. A case counts only when a digit is simply not a possibility: the ground truth lists no digit spelling for the word, the word itself is not a number, and no digit is accepted in the slot on either side — so the numeral cannot have spilled over from a real number standing nearby. Every one of the 6,028 is a digit the model introduced on its own.
The mechanism is mishearing toward the nearest number. Urdu بہتر, "better", sounds like bahattar — seventy-two — and comes back as 72. Urdu اسی, "that same", sounds like assī — eighty — and comes back as 80. Malayalam ഒരു, the indefinite article "a", comes back as 1. Speakers naming their own language, छत्तीसगढ़ी, get 36 — the chhattīs hiding inside the name. And this is not one language's quirk: the same substitution runs through Urdu, Bengali, Malayalam and beyond. The words are ordinary. The digits are invented.
Close to one English word in three comes back in a form nobody accepts.
A loanword is a word that belongs natively to English, however often everyday writing renders it in another script. The ground truth is generous with these: it carries optional spellings, so any accepted form of the word counts. Even so, across 187,358 spoken loanwords, Microsoft returns a form that exists nowhere in the lattice 29.03% of the time — close to one word in three. The rest of the field: 10.17–14.46%. These are not technical terms. They are "ma'am", "family", "design".
Each bar is one word's failures, split by what became of the word. Hover any segment.
| Word | Spoken | Microsoft wrong | Best comparator | What happens to it | Wrote instead |
|---|---|---|---|---|---|
मैम ma amHindi | 457 | 41.6% 190 of 457 | 4.2% | 19 168 19 · 168 · 3 | ऐम87ऍम54मम5माँ5अः3+14 more forms |
একচুয়ালি actuallyBengali | 182 | 64.8% 118 of 182 | 0.0% | 116 2 · 116 · 0 | কতুল্য88কোস্টিং1ফার্স্ট1একজনই1কিরে1+24 more forms |
ओके okayMarathi | 201 | 53.7% 108 of 201 | 1.5% | 79 22 79 · 22 · 7 | के4ओ3वा1स्ट्रेन1आपल्या1+12 more forms |
मॅम ma amMarathi | 114 | 86.0% 98 of 114 | 4.4% | 52 45 52 · 45 · 1 | मम3नाही2म2माम2मॅन2+32 more forms |
मॅडम madamMarathi | 231 | 42.4% 98 of 231 | 0.9% | 52 43 52 · 43 · 3 | मडम4†एक2नमस्कार1स्टुडंटे1om1+34 more forms |
ലാസ്റ്റ് lastMalayalam | 74 | 95.9% 71 of 74 | 2.7% | 8 63 8 · 63 · 0 | ലിസ്റ്4712ചേച്ചത്1ആസ്റ്1അവൻ1+11 more forms |
اوکے okayUrdu | 91 | 64.8% 59 of 91 | 5.5% | 32 24 32 · 24 · 3 | ہو3کہ3اک2کیا2اور2+12 more forms |
বেসিক্যালি basicallyBengali | 48 | 95.8% 46 of 48 | 2.1% | 5 41 5 · 41 · 0 | ব্যাসিক্যালয়28বেসিকালয়2বেসিকাল2†রেসিকলি1বিসকলে1+7 more forms |
নরমালি normallyBengali | 46 | 95.7% 44 of 46 | 0.0% | 4 40 4 · 40 · 0 | নর্মালয়321001এ1নরমেলিস1নর্মাল1†+4 more forms |
কোশ্চেন questionBengali | 41 | 100.0% 41 of 41 | 2.4% | 41 0 · 41 · 0 | কস্টিং24করছেন2বসেন2পাঞ্চার1তো1+11 more forms |
لائک likeUrdu | 39 | 100.0% 39 of 39 | 2.6% | 38 1 · 38 · 0 | لکے32ہے2مٹ1ایک1لگ1+1 more forms |
सॉरी sorryMarathi | 86 | 44.2% 38 of 86 | 0.0% | 23 15 23 · 15 · 0 | सारी4†नऊ1एक1अल्पकथा1अचं1+7 more forms |
کوشچن questionUrdu | 34 | 91.2% 31 of 34 | 2.9% | 3 28 3 · 28 · 0 | قوستیوں10کوسن2†نے2پہلے1پڑھی1+12 more forms |
شیئر shareUrdu | 47 | 61.7% 29 of 47 | 2.1% | 3 25 3 · 25 · 1 | سہارے20مجھے1سیر1سے1کیے1+1 more forms |
ایگزام examUrdu | 40 | 70.0% 28 of 40 | 0.0% | 3 25 3 · 25 · 0 | ےگزام8ےشَم2ےشام2ایکسان2جام1+10 more forms |
डिज़ाइन designHindi | 24 | 100.0% 24 of 24 | 0.0% | 24 0 · 24 · 0 | डिज़ैन21डिजैन्स1रिसाइन1मिनी1 |
സ്ട്രീറ്റ് streetMalayalam | 42 | 50.0% 21 of 42 | 0.0% | 20 1 · 20 · 0 | സ്ട്രീറ്8ട്രീറ്റ്2സ്പീഡ്1സ്വീറ്1നിങ്ങൾ1+7 more forms |
ബെസ്റ്റ് bestMalayalam | 20 | 100.0% 20 of 20 | 0.0% | 19 1 · 19 · 0 | ബേസ്ഡ്15ന്1ദിസ്1വെസ്റ്റ്1†ഫ്രണ്ട്1 |
डिजाइन designBhojpuri | 21 | 81.0% 17 of 21 | 0.0% | 17 0 · 17 · 0 | डिज़ैन15रहा1हो1 |
ப்ராடக்ட் productTamil | 19 | 63.2% 12 of 19 | 0.0% | 11 1 · 11 · 0 | ப்ரொடெக்ட்5மேம்4புராடக்ட்ஸ்1அ1 |
मैम ma amBhojpuri | 22 | 54.5% 12 of 22 | 0.0% | 11 0 · 11 · 1 | ऐम6ऍम5 |
எக்ஸாம்பிள் exampleTamil | 21 | 47.6% 10 of 21 | 0.0% | 1 9 1 · 9 · 0 | எக்ஸாம்7சொல்லி1அபார்ட்மெண்ட்1 |
அட்லீஸ்ட் at leastTamil | 22 | 36.4% 8 of 22 | 0.0% | 1 7 1 · 7 · 0 | லீஸ்ட்7 |
साइकिल cycleBhojpuri | 16 | 43.8% 7 of 16 | 0.0% | 1 6 1 · 6 · 0 | सैकल6 |
† this form is within one character of a spelling the reviewers accept, so it is more likely a gap in the reference than an error by the system. Words where such forms account for more than a quarter of the failures are not in this table at all.
The stacked bar is the point. On Marathi ओके, 79 of the 108 failures are not a misspelling — the word is simply not written. Same for मॅम and मॅडम, 53% each, and Urdu اوکے, 54%.
Chhattisgarhi has no row: no English loanword in its accepted-spelling lists clears the bar.
A tail of words it nearly always gets wrong.
2,310 distinct words are spoken at least ten times, wrong for Microsoft more than 40% of the time, and wrong for the best comparator less than 10% of the time. Together they account for 42,383 word errors. The failure is lexical rather than acoustic — the same word fails repeatedly, from different speakers, on audio the field transcribes.
Four, and that is the whole list. Once English loanwords move to Finding 03, there is almost no native Hindi word this model fails and the field does not.
| Word | Spoken | % of all words | Wrote instead | Microsoft wrong | × the median | Worst comparator | Sarvam Audio | Gemini 3 Pro | ElevenLabs Scribe v2 | Indic Conformer | Amazon Transcribe | Gemini 3 Flash |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
मिलेगी will get | 26 | 0.006% | मिलेंगेवोब | 100.0% | – | 3.8% | 0.0% | 0.0% | 0.0% | 0.0% | 3.8% | 0.0% |
यानी कि that is to say | 22 | 0.005% | घनिष्ठियानी कीयानी कीएफिशिएंट्ली यानी | 31.8% | 14.0× | 9.1% | 4.5% | 0.0% | 0.0% | 9.1% | 4.5% | 0.0% |
शहरी urban | 21 | 0.005% | शहरहैसभी | 23.8% | 3.3× | 14.3% | 4.8% | 14.3% | 4.8% | 0.0% | 9.5% | 10.0% |
घाट riverbank steps | 22 | 0.005% | भाटऔरआठ | 22.7% | 5.0× | 18.2% | 4.5% | 4.5% | 4.5% | 9.1% | 4.5% | 18.2% |
Some words simply never appear.
110,450 spoken words — 4.79% of the words the reviewers scored — appear nowhere in Microsoft's transcript for that clip: not in place, and not displaced elsewhere. The best comparator loses 1.24%, the worst 2.85%.
Audibility is not the explanation. Narrow the count to words Microsoft alone loses — words that at least five of the six comparators transcribed correctly from the same audio — and 57,140 remain. 52.1% of them are content words of four characters or more. Short function words are 26.9%. Three-character words are 19.1%.
Nor is it one word at a time. 6.2% of its deletion events remove three or more consecutive words, and those runs account for 19.8% of everything it loses.
A Microsoft-specific deletion is a reference word Microsoft omits entirely while at least five of the six comparators write it from the same audio. Runs and word classes are counted on that one base.
When a negation is lost, the sentence says the opposite.
A negation is the one word in a sentence that reverses it. Microsoft is wrong on 16.73% of the 45,008 negations spoken, against 4.87–11.44% for the field — but the split matters more than the total.
5.98% of negations are dropped: the word is not written anywhere in the transcript, and the sentence that comes out asserts what the speaker denied. The best comparator drops 1.03% — Microsoft loses a negation 5.8 times as often.
A further 10.76% are replaced by a different word, which damages the sentence in a way a reader cannot detect, because what comes out is fluent.
Each bar is the share of that word's spoken occurrences the system got wrong — either replaced with a different word or dropped entirely.
Signal conditions worsen and yet Microsoft moves less than anyone.
Every system in the field improves as recordings get cleaner and clips get longer. Microsoft included — 24.93% on the noisiest quartile falls to 17.01% on the cleanest. But laid over the six comparators, the *slope* of Microsoft's line is the shallowest of any system: a 1.48× improvement against a field range of 1.72×–2.42×. The same pattern repeats for utterance duration: 1.57× for Microsoft against 1.82×–2.95× for the field. Microsoft is the most stable system in the field — but at a floor the leaders never touch. Its cleanest quartile (17.01%) is worse than the weakest comparator's noisiest quartile (15.22%).
Any two systems in the release, district by district.
The district release scores every system on the same audio, district by district, in every language. Pick a language and a system for each map and the two render on one shared colour scale, so the darker map is the weaker system — Microsoft Speech-to-Text against Gemini 3 Pro by default, but any of the twenty-two systems against any other, in any of the fifteen languages. Hover a district to read both error rates and the gap between them.