We Benchmarked Our Speech Recognition Against Whisper in 59 Languages

We Benchmarked Our Speech Recognition Against Whisper in 59 Languages

6 min read

Most transcription tools tell you they support 100+ languages. Almost none tell you how well.

That is a convenient gap, because supporting a language and being any good at it are different things — and the difference is largest exactly where you can least afford it, in languages with little training data behind them.

So we measured. Here is the whole thing: every number, the method, and the languages where we lose.

---

The headline

Across 57 languages scored on both sides:

HarmarWhisper large-v3
Mean word error rate8.1%34.5%
Median word error rate4.1%14.6%
Languages won561

Whisper wins English. We win everything else that both systems can transcribe.

And there are two languages Whisper cannot attempt at all — it rejects Kyrgyz and Kinyarwanda as unsupported. We transcribe Kyrgyz at 5.8% and Kinyarwanda at 22.1%.

---

How this was measured

The data. FLEURS — Google's multilingual speech benchmark. It is the same sentences read aloud in every language, which is the property that makes a cross-language comparison meaningful at all: a language is not "harder" here just because its test set is harder. Twenty clips per language, sixty where we needed a larger sample. The baseline. OpenAI's `whisper-large-v3`, the model behind a large share of the captioning market, run on the identical audio files. The metric. Word error rate — insertions plus deletions plus substitutions, over the number of words in the reference. Lower is better. 5% means one word in twenty is wrong. The normalisation, which matters more than it sounds. Both transcripts are lowercased, stripped of punctuation, and have numbers put into one consistent form before scoring.

That last step is not cosmetic. We write "4 years"; the FLEURS reference says "four years", and so does Whisper. Without normalising, we get charged a full word error every time we format a number the way a subtitle should be formatted — on 14% of all clips. That single artefact was worth about a point of WER overall, and far more in some languages: Finnish 6.2% → 2.8%, Catalan 4.5% → 1.3%.

We normalise both sides. Scoring it the other way would have flattered Whisper for no reason other than that it spells numbers out.

---

Full results

Word error rate, lower is better. Both sides normalised identically.

LanguageHarmarWhisper large-v3Clips
Afrikaans9.1%35.6%20
Albanian8.8%56.0%60
Armenian9.3%38.7%20
Bashkir19.9%99.7%60
Basque6.2%40.3%60
Belarusian3.0%41.1%20
Bosnian6.4%16.0%20
Bulgarian2.3%9.6%20
Catalan4.1%5.5%20
Croatian1.9%10.1%20
Czech0.8%12.8%20
Danish6.2%12.7%20
Dutch3.1%5.8%20
English3.3%2.5%20
Estonian4.6%17.2%20
Finnish3.6%8.8%20
French1.8%3.4%20
Galician2.5%13.8%20
Georgian6.7%66.0%20
German2.9%3.1%20
Greek2.4%5.8%20
Hausa12.9%86.0%20
Hungarian3.9%14.6%20
Icelandic3.6%37.6%20
Indonesian3.0%6.7%20
Italian0.7%1.5%20
Kazakh4.0%39.3%20
Kinyarwanda22.1%not supported55
Kyrgyz5.8%not supported20
Latvian3.6%17.9%20
Lingala18.7%64.0%20
Lithuanian4.5%26.7%20
Luxembourgish23.8%86.3%20
Macedonian2.9%12.8%20
Malay2.9%8.9%20
Maltese5.4%68.7%20
Mongolian12.3%88.9%20
Norwegian4.1%8.1%20
Polish1.6%4.3%20
Portuguese2.0%3.0%20
Romanian1.9%8.1%20
Russian1.6%3.4%20
Serbian2.5%91.6%20
Shona15.6%113.0%20
Slovak0.5%6.4%20
Slovenian6.0%19.6%20
Somali40.5%95.1%20
Spanish3.2%4.3%20
Swahili5.3%30.7%20
Swedish5.5%9.3%20
Tagalog6.1%11.4%20
Tajik5.9%95.7%20
Tatar9.4%92.7%60
Turkish3.8%4.1%20
Turkmen44.6%99.3%60
Ukrainian4.3%10.5%20
Uzbek7.6%85.0%20
Vietnamese3.1%6.9%20
Yoruba75.2%100.2%20

---

Where the gap is largest

The pattern is not subtle: the fewer resources a language has, the wider the gap.

LanguageHarmarWhisper
Serbian2.5%91.6%
Uzbek7.6%85.0%
Tatar11.1%90.6%
Mongolian11.7%87.1%
Hausa12.9%86.0%
Maltese7.4%69.5%
Georgian6.7%65.1%
Albanian8.8%56.0%

At 85–92% word error rate, a transcript is not a rough draft. It is noise — there is nothing in it worth correcting, and you would be faster typing the whole thing yourself.

Meanwhile in French, German, Spanish and Italian, both systems are in low single digits and the practical difference is small. If you make videos in a major European language, most tools will serve you. The reason to care about this table is if you do not.

---

Where we lose, and where we are simply bad

A benchmark that only reports its wins is marketing. Ours:

English — Whisper wins. 2.5% against our 3.3%. English is the most over-served language in speech recognition and we are not going to beat a model trained overwhelmingly on it. German, Italian, Turkish — effectively a tie. We are ahead by a few tenths of a point, which is not a difference anyone experiences. Yoruba — we are bad. 75.2% as transcribed. Strip the diacritics before scoring and it falls to 19.7%, which tells you exactly where the failure is: we hear the words and we get the tone marks wrong. Yorùbá tone marks change meaning, so this is a real defect and not a formatting quibble. Whisper is worse still (100.2%), and that is no defence. We are fixing it, and until we have, we do not offer a Yoruba page. Turkmen (44.6%) and Somali (40.5%) are poor in absolute terms, even though Whisper is at 99.3% and 95.1%. Being twice as good as unusable is not the same as being good.

---

What this does not measure

Three limits, and the third is the one we care most about.

FLEURS is read speech. Someone reading a written sentence in a quiet room. Your videos are spontaneous, in a kitchen, with music. Every system scores better here than on real footage, ours included. Use these numbers to compare systems, not to predict what you will get. It is one voice at a time. No overlapping speakers, no crosstalk. It is monolingual — and that is the case we are actually built for. Every clip is one language, start to finish. Published research puts monolingual speech recognition 30–50% worse on code-switched audio, and Whisper's character error rate more than doubles, partly because it responds to a language switch by translating rather than transcribing.

Since mixing languages mid-sentence is how a great many people actually talk on camera — Spanglish, Taglish, Hinglish, Armenian with Russian and English — the monolingual number is not the number that matters most for us. We are running that benchmark separately and will publish it the same way: full table, method, and the losses.

---

Reproduce it

The dataset is public. The baseline model is public. The metric is standard, and the normalisation is described above. If you want to check any of this, you can — that is the point of naming everything.

If you make videos in a language the big tools treat as an afterthought, see how we handle yours.

Want subtitles for your video?

Try Harmar.ai for free. The AI turns your Armenian speech into trendy subtitles in a minute or two.

Try for free