Why Caption Tools Break When You Switch Languages Mid-Sentence

Why Caption Tools Break When You Switch Languages Mid-Sentence

7 min read

You film a video. You speak the way you actually speak — a sentence starts in one language, an English word lands in the middle, maybe a whole clause switches over and then switches back. Everyone around you talks like this. Nobody thinks about it.

Then you run it through an auto-caption tool, and the captions are wrong in a way that is hard to describe. Not misheard exactly. The words are there, but they have been quietly rewritten into one language. Or the borrowed words are missing. Or they have been spelled in an alphabet you would never have used.

This is not a bug in one app. It is a design decision that almost every captioning tool made, and it is worth understanding, because it tells you what to look for.

---

The single-language assumption

Nearly every auto-caption tool asks you one question before it starts: what language is this video in?

That question sounds harmless. It is actually the whole problem.

Once you answer it, the model commits. For the rest of the clip it is looking for words in that language. When you switch, the model does not stop and reconsider — it keeps going, and it does one of three things with the words that do not fit:

1. Translates them. You said an English word inside a Tagalog sentence; the transcript shows the Tagalog equivalent. You never said that.

2. Drops them. The word was not confidently in the selected language, so it did not survive.

3. Respells them. The borrowed word gets written in the target language's alphabet, producing something no one writes in real life.

All three produce a transcript that is plausible and wrong. That is the worst combination, because you have to read carefully to catch it — and you already know what you said, so your eye slides right over it.

---

Why this is so common

Because for most of the world's content, the assumption holds. A news broadcast is in one language. A lecture is in one language. The training data and the benchmarks that captioning models are measured against are overwhelmingly single-language, and a model that commits early is more accurate on that material.

The people it fails are the people who mix — which, in short-form video, is an enormous share of everyone:

  • Taglish in the Philippines — Tagalog and English, several times a sentence
  • Sheng in Kenya — Swahili and English, plus its own vocabulary
  • Hinglish in India — common enough that products exist just for it
  • Kazakh and Russian, Uzbek and Russian across Central Asia
  • Armenian with Russian and English, which is how most Armenian creators actually talk
  • Every diaspora community on earth, in every direction

If that is how you speak, the tool was not built for you. It was built for the news broadcast.

---

The alphabet problem nobody mentions

There is a second failure that only shows up in some languages, and it is worse because it looks like a typo rather than a mistake.

When a language is written in more than one alphabet — Kazakh and Uzbek are both written in Cyrillic and Latin in daily use — the tool has to pick one. And long videos are usually transcribed in several passes rather than all at once. If nothing pins the choice, different passes can pick different alphabets, and you get a transcript that is Cyrillic for four minutes and Latin for the next four. Same speaker. Same sentence structure. Two alphabets.

The related problem for borrowed words: if you say an English brand name inside a Kazakh sentence, should the transcript show it in Latin letters as everyone writes it, or transliterated into Cyrillic? There is a right answer, it is different per language, and a tool that never thought about the question will not be consistent about it.

---

What to check before you trust a tool

You do not need to understand any of the above to test it. Take a clip where you genuinely switch languages — not a clean one, a normal one — and look at four things:

1. Are the borrowed words still there? Search the transcript for a specific word you know you said in the other language. If it is missing or translated, stop there. 2. Is the spelling consistent? Find the same borrowed word in two places. It should be written identically both times. 3. Is the alphabet the same all the way down? Particularly on a video longer than a few minutes. Scroll to the end and compare it to the start. 4. Does the timing hold at the end? Many tools have a model estimate when each word was spoken rather than measuring it against the audio. Estimates drift, and the drift is worst at the end of a long clip — which is exactly where nobody checks.

If a tool passes those four on your real footage, it will probably hold up. If it fails the first one, nothing else matters.

---

What we do about it

Harmar was built for Armenian, where mixing with Russian and English is not an edge case — it is the default way people talk on camera. That constraint shaped the whole product, and it is why mixed speech is handled the way it is:

  • Words stay in the language they were spoken in. Nothing is translated into the surrounding language on the way to the transcript.
  • Spelling rules for borrowed words are written down per language, rather than left to whatever the model finds most probable that day.
  • The alphabet is fixed for the whole video for languages written in more than one, so a long clip cannot come back in two.
  • Timing is measured against the audio, word by word, instead of being estimated — so it does not drift as the video runs.

And the editor exists because none of this is ever perfect. Every word is editable with the video playing beside it, and every word's timing can be dragged by hand. Where the timing is an estimate rather than a measurement, it says so instead of pretending to precision.

See which languages we handle →

---

The short version

If you speak one language on camera, almost any captioning tool will do.

If you switch — and most people making short-form video do — then the question to ask is not "does this tool support my language?" Nearly all of them will say yes. The question is what does it do with the words that are not in that language, and the only way to find out is to run your own footage through it and read the result carefully.

Want subtitles for your video?

Try Harmar.ai for free. The AI turns your Armenian speech into trendy subtitles in a minute or two.

Try for free