BENCHMARK REPORT
Release
August 2026
Series
Voice of India
Model
Microsoft STT

How Well Does Microsoft Speech-to-Text Understand Indian Speech?

Results from the Voice of India Benchmark — testing Microsoft Speech-to-Text across 8 of the benchmark's 15 languages.

COVERAGE

Microsoft is scored on eight of fifteen languages.

Microsoft Speech-to-Text does not support these seven languages:

  • Assamese
  • Gujarati
  • Kannada
  • Maithili
  • Odia
  • Punjabi
  • Telugu

Every figure in this report is confined to the eight it does cover.

Microsoft is scored on these eightMicrosoft does not support these seven
SystemHindiBengaliMarathiTamilMalayalamBhojpuriChhattisgarhiUrduAssameseGujaratiKannadaMaithiliOdiaPunjabiTelugu
Indic Transcribe Core3.54.35.79.011.513.313.65.08.29.27.411.38.78.311.4
Saaras v33.85.26.59.112.217.914.07.59.19.78.814.211.18.613.5
Gemini 3 Pro4.76.78.811.716.315.313.46.917.013.414.020.317.912.718.4
ElevenLabs Scribe v25.97.910.315.316.417.614.219.211.818.013.917.713.219.6
IndicConformer6.18.910.716.020.930.124.47.410.815.516.015.412.212.819.8
Gemini 3 Flash7.011.014.117.323.020.121.610.924.920.218.427.923.217.825.4
Microsoft Streaming STT7.913.626.922.133.431.825.617.117.2
Microsoft STT7.523.329.322.835.731.626.223.3
Gemma E4B9.119.624.437.944.427.024.014.245.027.331.036.744.423.441.6
GPT Realtime10.315.820.523.637.631.227.339.726.232.930.0
OmniASR LLM 7B v29.620.924.540.648.826.320.714.823.931.935.044.672.331.848.6
Amazon Transcribe4.86.57.712.320.627.522.413.810.913.812.313.9
Deepgram Nova 39.327.142.166.439.734.051.041.1
Ringg12.0
Systems marked ⁺ are averaged over the languages they cover among the eight, shown in brackets.OI-WER as published in the Voice of India benchmark report. Lower is better.Ranked on the eight languages Microsoft covers, Microsoft Speech-to-Text is tenth of seventeen — 29.4 against 12.4 for the best system on the same eight.Microsoft's own streaming product scores 28.3 on the same eight languages, ahead of the non-streaming one measured here.

Findings across eight languages

TEN KEY FINDINGS

Our re-measurement, eight languages

LanguageMicrosoft STTSaaras v3Gemini 3 ProElevenLabs Scribe v2IndicConformerAmazon TranscribeGemini 3 Flash
Hindi7.53.84.75.96.14.87.0
Bengali23.35.26.77.98.96.511.0
Marathi29.36.58.810.310.77.714.1
Tamil22.89.111.715.316.012.317.3
Malayalam35.712.216.316.420.920.623.0
Bhojpuri31.617.915.317.630.127.520.1
Chhattisgarhi26.214.013.414.224.422.421.6
Urdu23.37.56.919.27.410.9
Complete 8-language comparison across the field. Word error rate against the accepted-spelling lattice, on identical audio. Amazon Transcribe is not scored in Urdu.
Finding 01 · SIGNATURE · INVERSE TEXT NORMALIZATION

When nobody said a time, it writes a clock anyway.

A clock-time opportunity is a clip whose reference contains no time word and no clock — a clip in which nobody said a time. There are 154,555 of them across the eight languages this system covers. It writes a colon-formatted clock on 3,290 of them. The six scored comparators, on the same clips, write one 0 times.

It is not a mis-hearing. A number word is present in the audio — “two”, “a”, “a minute” — and a post-processing stage decides that a number standing near another number must be a time, and formats it as one. The minutes it prints were never spoken and appear in no accepted spelling of the reference.

Clips with a fabricated clock
3,290
+46 more where a time was said but the clock is wrong
Share of opportunities
2.10%
The entire field
0
six systems, whole corpus
Reduplicated phrase
12 of 12
field 1–3
Fabricated clock times — every system in the release
The same 154,555 clips in which nobody said a time. Each dot is one system, placed by how many of those clips it wrote a colon-clock on anyway.

Microsoft writes 3,244. The next highest is 20 — 124× lower.

15 of the 23 systems never do it once. Microsoft's own streaming product is one of them: it fabricates a clock zero times on the same clips.

Ten of them, in eight languages
Spoken phrase, then the clock it printed.
दो दो बजे2:02Hindi
"two, two o'clock" — the speaker repeats the hour
सात आठ बजे7:08Hindi
"seven, eight o'clock" — two candidate hours, not a minute
দুটো কল2:00Bengali
"two calls" — no hour, no clock, no time word
একটা গ্রামে1:00Bengali
"a village" — the indefinite article
दोन झाडांना2:00Marathi
"two plants" — nothing about time in the clip
ஒரு நிமிடம்1:00Tamil
"a minute" — a duration, not a clock reading
ഒരു ഒരു മണിക്കൂർ1:01Malayalam
"an hour" — the repeated article becomes minutes
दू दु बजे2:02Bhojpuri
"two, two o'clock" — the hour repeated
एक दो बजे1:02Chhattisgarhi
"one or two o'clock" — an approximation
بارہ تیرہ گھنٹہ12:13Urdu
"twelve, thirteen hours" — a length of time
Microsoft Speech-to-Text
Finding 02 · SIGNATURE · INVERSE TEXT NORMALIZATION

Ordinary words come back as digits.

Microsoft's model keeps returning digits in sentences where no number exists. These are not misread numbers — they are ordinary words, replaced by numerals. On this test it happens 6,028 times. Across the six comparators combined, it happens 4 times.

"No number exists" is meant literally. A case counts only when a digit is simply not a possibility: the ground truth lists no digit spelling for the word, the word itself is not a number, and no digit is accepted in the slot on either side — so the numeral cannot have spilled over from a real number standing nearby. Every one of the 6,028 is a digit the model introduced on its own.

The mechanism is mishearing toward the nearest number. Urdu بہتر, "better", sounds like bahattar — seventy-two — and comes back as 72. Urdu اسی, "that same", sounds like assī — eighty — and comes back as 80. Malayalam ഒരു, the indefinite article "a", comes back as 1. Speakers naming their own language, छत्तीसगढ़ी, get 36 — the chhattīs hiding inside the name. And this is not one language's quirk: the same substitution runs through Urdu, Bengali, Malayalam and beyond. The words are ordinary. The digits are invented.

Microsoft
6,028
The entire field
4
Urdu اسی
95.1%
field 2.5–8.9%
Bengali একটা
92.2%
field 0.8–3.5%
The four words the model turns into numbers
Where each comparator lands on the four words, and how far Microsoft diverges from the field.
اسی"that same" · Urdu80بہتر"better" · Urdu72একটা"a / one" · Bengali00ഒരു"a" · Malayalam10255075100%
Words that were turned into numbers
None of these is a number. No digit was spoken, and no digit is an accepted spelling in any of these slots.
ഒരു
"a"
Malayalam1624 times
اسی
"that same"
Urdu80159 times
ஒரு
"a / one"
Tamil11143 times
بہتر
"better"
Urdu7283 times
ഒരു
"a"
Malayalam1171 times
നല്ലൊരു
"a good"
Malayalamനല്ല 150 times
ഓരോ
"each"
Malayalam131 times
വലിയൊരു
"a big"
Malayalamവലിയ 131 times
"that"
Malayalam125 times
শো
"show"
Bengali10018 times
বার
"times"
Bengali1217 times
ஒரு ஒரு
"one by one"
Tamil1115 times
നല്ലൊരു
"a good"
Malayalam114 times
വലിയൊരു
"a big"
Malayalam110 times
একটানা
"continuously"
Bengali00 না7 times
ہو
"is"
Urdu1005 times
Microsoft Speech-to-Text
Finding 03 · LEXICAL · ENGLISH LOANWORDS

Close to one English word in three comes back in a form nobody accepts.

A loanword is a word that belongs natively to English, however often everyday writing renders it in another script. The ground truth is generous with these: it carries optional spellings, so any accepted form of the word counts. Even so, across 187,358 spoken loanwords, Microsoft returns a form that exists nowhere in the lattice 29.03% of the time — close to one word in three. The rest of the field: 10.17–14.46%. These are not technical terms. They are "ma'am", "family", "design".

Loanwords spoken
184,113
on a common base
Microsoft wrong
28.85%
Field
10.09–14.45%
Multiple
2.4×
English loanwords, every system
Share of English word tokens written in a form no reviewer accepts. Every system is scored on the same words.
Microsoft Speech-to-Text28.85%53,120 of 184,113Indic Conformer14.45%26,604 of 184,113Amazon Transcribe12.85%21,513 of 167,464Gemini 3 Flash12.80%23,559 of 184,113ElevenLabs Scribe v211.20%20,612 of 184,113Gemini 3 Pro10.10%18,594 of 184,113Sarvam Audio10.09%18,568 of 184,113012.52537.550%
‡ Amazon Transcribe returns nothing in Urdu, so no Urdu clip has output from all seven systems. It is scored on the seven-system common base of 167,464 tokens, covering the other seven languages.
Every system is scored on the same loanword tokens: the 184,113 tokens in the 153,482 clips that all six of these systems transcribed. Amazon Transcribe returns nothing in Urdu, so no Urdu clip has output from all seven; it is scored on the seven-system common base of 167,464 tokens instead, which covers the other seven languages.
Per-system denominators would have scored Amazon Transcribe on 170,468 tokens and Gemini 3 Pro on 187,358 — a 10% difference in how much speech each was asked to handle.
English loanwords, all languages
Every word here is an English loanword the reviewers list with its English spelling. Microsoft is wrong on at least a third of the times it is spoken and at least three times as often as the best comparator. Loanwords appear only here — Finding 04 excludes them.
droppedwritten as something elsedisplaced

Each bar is one word's failures, split by what became of the word. Hover any segment.

WordSpokenMicrosoft wrongBest comparatorWhat happens to itWrote instead
मैम
ma amHindi
45741.6%
190 of 457
4.2%
19
168
19 · 168 · 3
ऐम87ऍम54मम5माँ5अः3+14 more forms
একচুয়ালি
actuallyBengali
18264.8%
118 of 182
0.0%
116
2 · 116 · 0
কতুল্য88কোস্টিং1ফার্স্ট1একজনই1কিরে1+24 more forms
ओके
okayMarathi
20153.7%
108 of 201
1.5%
79
22
79 · 22 · 7
के43वा1स्ट्रेन1आपल्या1+12 more forms
मॅम
ma amMarathi
11486.0%
98 of 114
4.4%
52
45
52 · 45 · 1
मम3नाही22माम2मॅन2+32 more forms
मॅडम
madamMarathi
23142.4%
98 of 231
0.9%
52
43
52 · 43 · 3
मडम4एक2नमस्कार1स्टुडंटे1om1+34 more forms
ലാസ്റ്റ്
lastMalayalam
7495.9%
71 of 74
2.7%
8
63
8 · 63 · 0
ലിസ്റ്4712ചേച്ചത്1ആസ്റ്1അവൻ1+11 more forms
اوکے
okayUrdu
9164.8%
59 of 91
5.5%
32
24
32 · 24 · 3
ہو3کہ3اک2کیا2اور2+12 more forms
বেসিক্যালি
basicallyBengali
4895.8%
46 of 48
2.1%
5
41
5 · 41 · 0
ব্যাসিক্যালয়28বেসিকালয়2বেসিকাল2রেসিকলি1বিসকলে1+7 more forms
নরমালি
normallyBengali
4695.7%
44 of 46
0.0%
4
40
4 · 40 · 0
নর্মালয়3210011নরমেলিস1নর্মাল1+4 more forms
কোশ্চেন
questionBengali
41100.0%
41 of 41
2.4%
41
0 · 41 · 0
কস্টিং24করছেন2বসেন2পাঞ্চার1তো1+11 more forms
لائک
likeUrdu
39100.0%
39 of 39
2.6%
38
1 · 38 · 0
لکے32ہے2مٹ1ایک1لگ1+1 more forms
सॉरी
sorryMarathi
8644.2%
38 of 86
0.0%
23
15
23 · 15 · 0
सारी4नऊ1एक1अल्पकथा1अचं1+7 more forms
کوشچن
questionUrdu
3491.2%
31 of 34
2.9%
3
28
3 · 28 · 0
قوستیوں10کوسن2نے2پہلے1پڑھی1+12 more forms
شیئر
shareUrdu
4761.7%
29 of 47
2.1%
3
25
3 · 25 · 1
سہارے20مجھے1سیر1سے1کیے1+1 more forms
ایگزام
examUrdu
4070.0%
28 of 40
0.0%
3
25
3 · 25 · 0
ےگزام8ےشَم2ےشام2ایکسان2جام1+10 more forms
डिज़ाइन
designHindi
24100.0%
24 of 24
0.0%
24
0 · 24 · 0
डिज़ैन21डिजैन्स1रिसाइन1मिनी1
സ്ട്രീറ്റ്
streetMalayalam
4250.0%
21 of 42
0.0%
20
1 · 20 · 0
സ്ട്രീറ്8ട്രീറ്റ്2സ്പീഡ്1സ്വീറ്1നിങ്ങൾ1+7 more forms
ബെസ്റ്റ്
bestMalayalam
20100.0%
20 of 20
0.0%
19
1 · 19 · 0
ബേസ്ഡ്15ന്1ദിസ്1വെസ്റ്റ്1ഫ്രണ്ട്1
डिजाइन
designBhojpuri
2181.0%
17 of 21
0.0%
17
0 · 17 · 0
डिज़ैन15रहा1हो1
ப்ராடக்ட்
productTamil
1963.2%
12 of 19
0.0%
11
1 · 11 · 0
ப்ரொடெக்ட்5மேம்4புராடக்ட்ஸ்11
मैम
ma amBhojpuri
2254.5%
12 of 22
0.0%
11
0 · 11 · 1
ऐम6ऍम5
எக்ஸாம்பிள்
exampleTamil
2147.6%
10 of 21
0.0%
1
9
1 · 9 · 0
எக்ஸாம்7சொல்லி1அபார்ட்மெண்ட்1
அட்லீஸ்ட்
at leastTamil
2236.4%
8 of 22
0.0%
1
7
1 · 7 · 0
லீஸ்ட்7
साइकिल
cycleBhojpuri
1643.8%
7 of 16
0.0%
1
6
1 · 6 · 0
सैकल6

† this form is within one character of a spelling the reviewers accept, so it is more likely a gap in the reference than an error by the system. Words where such forms account for more than a quarter of the failures are not in this table at all.

The stacked bar is the point. On Marathi ओके, 79 of the 108 failures are not a misspelling — the word is simply not written. Same for मॅम and मॅडम, 53% each, and Urdu اوکے, 54%.

Chhattisgarhi has no row: no English loanword in its accepted-spelling lists clears the bar.

Microsoft Speech-to-Text
Finding 04 · LEXICAL · THE LONG TAIL

A tail of words it nearly always gets wrong.

2,310 distinct words are spoken at least ten times, wrong for Microsoft more than 40% of the time, and wrong for the best comparator less than 10% of the time. Together they account for 42,383 word errors. The failure is lexical rather than acoustic — the same word fails repeatedly, from different speakers, on audio the field transcribes.

Words in the tail
2,310
Word errors
42,383
Microsoft threshold
>40%
Best comparator
<10%
The native words it gets wrong that the others do not
A word qualifies only if Microsoft fails it at least two and a half times as often as the median comparator, at least fifteen points higher in absolute terms, and no comparator fails it more than a quarter of the time. That last test is new: if another system is also badly wrong on a word, the problem is more likely the recording or the reference than this model.

Four, and that is the whole list. Once English loanwords move to Finding 03, there is almost no native Hindi word this model fails and the field does not.

WordSpoken% of all wordsWrote insteadMicrosoft wrong× the medianWorst comparatorSarvam AudioGemini 3 ProElevenLabs Scribe v2Indic ConformerAmazon TranscribeGemini 3 Flash
मिलेगी
will get
260.006%
मिलेंगेवोब
100.0%3.8%0.0%0.0%0.0%0.0%3.8%0.0%
यानी कि
that is to say
220.005%
घनिष्ठियानी कीयानी कीएफिशिएंट्ली यानी
31.8%14.0×9.1%4.5%0.0%0.0%9.1%4.5%0.0%
शहरी
urban
210.005%
शहरहैसभी
23.8%3.3×14.3%4.8%14.3%4.8%0.0%9.5%10.0%
घाट
riverbank steps
220.005%
भाटऔरआठ
22.7%5.0×18.2%4.5%4.5%4.5%9.1%4.5%18.2%
Microsoft Speech-to-Text
Finding 05 · DELETION · WORDS THAT NEVER APPEAR

Some words simply never appear.

110,450 spoken words — 4.79% of the words the reviewers scored — appear nowhere in Microsoft's transcript for that clip: not in place, and not displaced elsewhere. The best comparator loses 1.24%, the worst 2.85%.

Audibility is not the explanation. Narrow the count to words Microsoft alone loses — words that at least five of the six comparators transcribed correctly from the same audio — and 57,140 remain. 52.1% of them are content words of four characters or more. Short function words are 26.9%. Three-character words are 19.1%.

Nor is it one word at a time. 6.2% of its deletion events remove three or more consecutive words, and those runs account for 19.8% of everything it loses.

Words lost
110,450
Share of all speech
4.79%
Field
1.24–2.85%
Multiple
2.5×
Words that never appear, every system
Deleted words out of every hundred spoken, system by system.
Microsoft Speech-to-Text4.79 in every 100 wordsGemini 3 Pro1.24 in every 100 wordsSarvam Audio1.25 in every 100 wordsElevenLabs Scribe v21.48 in every 100 wordsAll seven systemsMicrosoft Speech-to-Text4.79%Indic Conformer2.85%Gemini 3 Flash2.36%Amazon Transcribe2.33%ElevenLabs Scribe v21.48%Sarvam Audio1.25%Gemini 3 Pro1.24%01.32.53.85%
What it drops, when nobody else does
Microsoft-specific deletions only: words at least five of the six comparators transcribe from the same audio.
content word (4 characters or more)52.1% · 29,776
short function word (1-2 characters)26.9% · 15,356
three-character word19.1% · 10,901
negation1.9% · 1,107
Words removed at once
1 word80.7%
2 words13.1%
3+ words6.2%

A Microsoft-specific deletion is a reference word Microsoft omits entirely while at least five of the six comparators write it from the same audio. Runs and word classes are counted on that one base.

Deletions, every system
Reference words omitted entirely, on identical audio.
Microsoft Speech-to-Text158,605Sarvam Audio46,682Gemini 3 Pro48,900ElevenLabs Scribe v255,151Indic Conformer103,634Amazon Transcribe76,319Gemini 3 Flash87,291
Microsoft Speech-to-Text
Finding 06 · SEMANTIC INVERSION · NEGATION

When a negation is lost, the sentence says the opposite.

A negation is the one word in a sentence that reverses it. Microsoft is wrong on 16.73% of the 45,008 negations spoken, against 4.87–11.44% for the field — but the split matters more than the total.

5.98% of negations are dropped: the word is not written anywhere in the transcript, and the sentence that comes out asserts what the speaker denied. The best comparator drops 1.03% — Microsoft loses a negation 5.8 times as often.

A further 10.76% are replaced by a different word, which damages the sentence in a way a reader cannot detect, because what comes out is fluent.

Negations spoken
45,008
Dropped entirely
5.98%
Replaced
10.76%
Vs best on drops
5.8×
Dropped, or replaced
Share of spoken negations, worst first.
Dropped — the negation is not written at allReplaced — a different word appears in its place
Microsoft Speech-to-Text5.98% + 10.76%Gemini 3 Flash3.08% + 8.36%Indic Conformer3.94% + 7.21%Amazon Transcribe2.22% + 6.21%ElevenLabs Scribe v21.52% + 5.44%Gemini 3 Pro1.13% + 4.79%Sarvam Audio1.03% + 3.85%% of spoken negations
What the sentence becomes
Four clips where the negation is lost or replaced.
Hindiनहीं
Reference
उनके ध्यान थोड़ा भी नहीं दे पाते है
Microsoft
उन पे ध्यान थोड़ा भी दे पाता है।
Means: they cannot give them even a little attention
Reads as: they can give them a little attention
droppedclip 31913
Urduنہ
Reference
تاکہ کال کو ڈسٹربنس نہ ہو
Microsoft
کال کو ریسرن ہو۔
Means: so that the call is not disturbed
Reads as: so that the call is disturbed
droppedclip 128453
Chhattisgarhiनी
Reference
ता हमहूं कबू नी जा रे बाबू
Microsoft
मू मन कबू निजा रे बाबू
Means: I never go there either
Reads as: the negation is absorbed into the verb and no longer reads as one
droppedclip 332974
Malayalamഅല്ല
Reference
അല്ല ഒരു എഴുത്തുകാരൻ വന്നിട്ട്
Microsoft
1 എഴുത്തുകാരം വന്നിട്ട്
Means: no — a writer comes along
Reads as: 1 writer comes along
substitutedclip 237930
Which negations fail, within each language
Bars are compared only against the best comparator on the same word in the same language.

Each bar is the share of that word's spoken occurrences the system got wrong — either replaced with a different word or dropped entirely.

Microsoft — % of occurrences wrongBest comparator on the same word — % wrong
Regional negativesStandard negatives66.8% against 1.50% — a factor of 45नइखेBhojpuri232 spokenनाMarathi2,203 spokenBhojpuri367 spokenनाहीMarathi3,524 spokenनाहीयेMarathi317 spokenनाBhojpuri2,094 spokenഇല്ലMalayalam337 spokenഅല്ലMalayalam497 spokenنہUrdu365 spoken0255075100%
% of occurrences wrong
Microsoft Speech-to-Text
Finding 07 · CONTROL · AUDIO QUALITY, SPEAKING RATE, DURATION

Signal conditions worsen and yet Microsoft moves less than anyone.

Every system in the field improves as recordings get cleaner and clips get longer. Microsoft included — 24.93% on the noisiest quartile falls to 17.01% on the cleanest. But laid over the six comparators, the *slope* of Microsoft's line is the shallowest of any system: a 1.48× improvement against a field range of 1.72×–2.42×. The same pattern repeats for utterance duration: 1.57× for Microsoft against 1.82×–2.95× for the field. Microsoft is the most stable system in the field — but at a floor the leaders never touch. Its cleanest quartile (17.01%) is worse than the weakest comparator's noisiest quartile (15.22%).

Microsoft slope (noise)
1.48×
field 1.72×–2.42×
Microsoft slope (duration)
1.57×
field 1.82×–2.95×
Stability rank
1 of 7
flattest curve on both axes
Floor vs field ceiling
17.01% > 15.22%
MS cleanest worse than field's worst
Microsoft has the flattest curve in the field
Per-model OI-WER across four quartiles. Microsoft's slope from worst to best conditions is the shallowest of any comparator — a stable high floor rather than a competitive ceiling.
By audio quality (DNSMOS quartile)
Microsoft slope 1.48×
05101520Q1 NoisyQ2 Mild-noiseQ3 Mid-qualityQ4 CleanMicrosoft STTSarvam AudioSaarika 2.5Amazon TranscribeGemini 3 ProIndicConformerElevenLabs Scribe v2
OI-WER (%)
By utterance duration
Microsoft slope 1.57×
05101520Q1 Short (≤3.8s)Q2 (3.8–6.7s)Q3 (6.7–12s)Q4 Long (>12s)Microsoft STTSarvam AudioSaarika 2.5Amazon TranscribeGemini 3 ProIndicConformerElevenLabs Scribe v2
OI-WER (%)
Microsoft STTAmazon TranscribeElevenLabs Scribe v2Gemini 3 ProIndicConformerSaarika 2.5Sarvam Audio
Hindi subset — 22,060 clips. Aggregate 8-language numbers appear in the metric strip and in the field-band chart above; per-comparator slopes are Hindi-only because that is where the audit was run.
Finding 08 · GEOGRAPHY · SYSTEM VERSUS SYSTEM

Any two systems in the release, district by district.

The district release scores every system on the same audio, district by district, in every language. Pick a language and a system for each map and the two render on one shared colour scale, so the darker map is the weaker system — Microsoft Speech-to-Text against Gemini 3 Pro by default, but any of the twenty-two systems against any other, in any of the fifteen languages. Hover a district to read both error rates and the gap between them.

Languages on the map
15
Systems comparable
22
Microsoft, overall
28.0%
Best in the field
13.0%
Sarvam-Omni
Every district, two systems at a time
Pick a language, then any two systems in the release, and read them district by district on identical audio. Only the districts that carry scored clips in the chosen language can be shaded; the rest of the country has no audio for it in the benchmark. Both maps share one colour scale, so the darker map is the weaker system.
Language
System A
System B
microsoft_STT
Loading districts…
gemini-3-pro
Loading districts…
Word error rate, % — one shared scale on both maps
01020+
no scored audio