Contents
AI subtitle tools ranked by accuracy
Speech recognition accuracy is measured with WER — the share of words the model got wrong, missed or added that were never said. To make the number tangible, let me translate it into work: an average speaker produces around 140 words per minute. So 2% WER is roughly three errors for every minute of video, and 10% is already fourteen.
The Artificial Analysis ranking is built on eight hours of audio from three sets: voice assistant conversations, European Parliament speeches and financial earnings calls. As of September 2026:
| Recognition engine | Words in error (AA-WER) | Approx. fixes per minute of video | Price per 1,000 min |
|---|---|---|---|
| Microsoft MAI-Transcribe-2 | 2.0% | ~3 | $1.67 |
| ElevenLabs Scribe v2 | 2.2% | ~3 | $3.67 |
| Google Gemini 3.5 Transcribe | 2.6% | ~4 | $5.00 |
| OpenAI GPT Transcribe | 3.3% | ~5 | $4.50 |
| AssemblyAI Universal | 3.8% | ~5 | $2.50 |
| OpenAI GPT-4o Transcribe | 4.0% | ~6 | $6.00 |
| Whisper Large v3 (fal.ai) | 4.1% | ~6 | $1.15 |
| Speechmatics Melia | 4.9% | ~7 | $4.00 |
| Deepgram Nova-3 | 5.2% | ~7 | $4.30 |
| Whisper Large v3 (Replicate) | 10.1% | ~14 | $4.23 |
One caveat that rarely gets mentioned: these datasets are predominantly English. The leader in English is not necessarily the leader in Ukrainian, and for our language the order in the table may look different.
The top of this ranking shifts almost monthly, so it is worth checking the current version before you choose.
Why the same model delivers different accuracy
Look at the two Whisper Large v3 rows. Same model, yet 4.1% errors at one provider and 10.1% at another — more than a twofold difference.
The reason is implementation: how the audio is split into chunks, how silence is handled, which decoding settings are used. There is a practical takeaway in this for you. A “powered by Whisper” line in a service description guarantees nothing. Two apps built on the same model can produce completely different results, so test the tool itself rather than the model under the hood.

Which service recognises Ukrainian most accurately for subtitles
There is unfortunately no independent ranking for Ukrainian specifically. That leaves developer benchmarks and your own testing.
- ElevenLabs Scribe — the best published figure: 3.1% errors on FLEURS. The developer places Ukrainian in the group of languages with “excellent accuracy” (under 5% errors).
- YouTube auto-captions have supported Ukrainian since 2022 and are free for your own videos.
- Instagram launched auto-captions with 17 languages, and Ukrainian was not among them. Check whether it has appeared in your account before counting on it.
- CapCut, TikTok and video editors — the list of auto-caption languages depends on the app version, so check in the settings rather than in the app store description.
Now for a trap worth knowing about. FLEURS is a dataset where people read prepared sentences aloud: clearly, without noise, without slang. Real video with surzhyk, English terms and background music will produce noticeably more than 3.1% errors on any tool.
That is why my working method is simple: I take 2–3 minutes of an actual client video, run it through two or three services and count how many words need fixing. It takes half an hour and gives a more accurate answer than any ranking.
If what you need is a transcript of calls and meetings rather than video, the tools and the criteria are different, and there is a separate article about them: “AI tools for transcribing calls”.
The best AI subtitle generator: what to pick for each task
There is no single best tool “for everything”. There is a best tool for a specific task:
| Task | What to choose | Why |
|---|---|---|
| Ukrainian-language video for YouTube | YouTube auto-captions + editing in Studio | free, Ukrainian supported, separate file right away |
| Burned-in subtitles for Reels and TikTok | a video editor with auto-captions | styling, word animation, right in the frame |
| Maximum accuracy, large volumes | engines from the table via API: Scribe, Gemini, AssemblyAI | top of the ranking, pay per minute |
| Free, unlimited, confidential | Whisper via Subtitle Edit on your own computer | the file never leaves your device |
| Free online, a few videos a day | TurboScribe | nothing to install |
| Interviews and podcasts with several speakers | an engine with speaker diarisation | it labels who is speaking |
For a team producing content for several platforms at once, the most convenient setup is this: generate the subtitles once in an accurate service, proofread them, save them as an SRT file and import that file into the editor for each platform. That way the edits are made once instead of three times.
The best free subtitle generators
| Tool | What it runs on | Limit | Ukrainian | Export |
|---|---|---|---|---|
| YouTube Studio | Google recognition | no limit for your own videos | yes | downloadable subtitle file |
| Subtitle Edit | Whisper, locally | no limit | yes | SRT and other formats |
| TurboScribe | Whisper Large v3 | 3 files a day, up to 30 min each | yes | SRT, TXT, DOCX, PDF |
Each of the three comes with its own trade-off:
- YouTube — the simplest, but it only works with videos you have uploaded to your channel. For Reels you will first have to upload the clip to YouTube as private.
- Subtitle Edit — the most freedom: open-source software, Whisper models download automatically on first run, and the video is never sent anywhere. The downside is that speed depends on your computer; on a weak laptop an hour-long video can take a long time to process.
- TurboScribe — the most convenient online option, but a limit of three files a day suits a blogger, not an agency.
Why AI makes mistakes in subtitles
Recognition errors are not random. They cluster in a few predictable places, and once you know those places, proofreading subtitles gets twice as fast.

Hallucinations on pauses and music
The nastiest kind of error is when the model writes a phrase that was never in the audio at all. Researchers from several US universities tested Whisper and found that roughly 1% of transcriptions contained entirely invented phrases or sentences. What is more, 38% of those inventions were not neutral but harmful: aggressive language, false statements or fabricated appeals to authority. Hallucinations occurred most often where the speech contained long pauses.
The practical takeaway:
- Generate subtitles before you add background music. Music under a quiet voice, plus pauses, is the most fertile ground for invented text.
- Proofread the passages after pauses more carefully than the rest of the text: that is where the model most often fills in what it did not hear.
Proper names, terminology and numbers
A brand name, a guest’s surname, a technical term — the model will substitute the closest word it knows. Numbers it may write out in words or as digits, and inconsistently within the same video.
Some services let you supply a list of specific words or a prompt with context in advance. If you regularly shoot content about one product or industry, this is the most effective way to cut down on edits: five minutes spent on a glossary saves dozens of corrections in every video.
Can AI recognise several languages in one video
Poorly, and for Ukrainian content this is one of the most common problems.
Whisper, which powers a large share of the free tools, detects the language at the start of the recording and then assigns a single language to each 30-second chunk. If a video starts in Ukrainian and the speaker switches to English halfway through, the model may write the English phrases in Cyrillic letters or translate them. Surzhyk and English terms inside a Ukrainian sentence break recognition even harder: there the language changes several times within one phrase.
What helps:
- Set the language manually rather than relying on auto-detection, especially if the first seconds of the video are music, an intro or another language.
- Split the video into parts by language if it contains long segments in different languages.
- Budget more proofreading time for videos where English terms appear inside Ukrainian sentences. That kind of mixing remains difficult for every model.
Frequently asked questions (FAQ)
What determines the quality of automatic transcription?
Above all, the sound. A lavalier mic 15–20 cm from the mouth, a separate audio track without music and a room without echo deliver a bigger accuracy gain than switching to a more expensive service.
Why does no AI deliver 100% subtitle accuracy?
Because conversational speech is ambiguous even for humans: professional transcribers get about 5% of words wrong on phone conversations. Unclear pronunciation, overlapping voices and identical-sounding words have no single correct answer.
How do I choose a subtitle tool for a specific platform (YouTube, Instagram, TikTok)?
For YouTube, a separate subtitle file generated in Studio or uploaded from another service. For Instagram and TikTok, burned-in subtitles from a video editor, because feeds are often scrolled with the sound off and the text has to be in the frame itself.
Burned-in subtitles or a separate file — which is better?
A separate file can be switched off by the viewer, corrected without re-editing and extended with translations. Burned-in subtitles are always visible and can be styled to match the brand, but any error in them requires a new render of the video.
Which format should I save subtitles in: SRT or VTT?
SRT if you need maximum compatibility: almost every editor and platform opens it. VTT is a web format that additionally supports styling and text positioning on screen.
How many characters should a subtitle line have?
The streaming platform standard is up to 42 characters per line and no more than two lines on screen. For vertical video the lines are made shorter, because the frame is narrower.
Can subtitles be translated into another language automatically?
Yes, YouTube translates subtitles automatically and most services have a translation feature. The quality will be lower than the original transcription, because recognition errors stack on top of translation errors, so the source text is worth proofreading before translation.
Is it safe to upload video to online subtitle services?
For public content, yes. For recordings containing trade secrets or personal data it is better to run Whisper locally through Subtitle Edit: the file never leaves your computer.
Does AI punctuate and capitalise correctly?
Not always, and accuracy rankings do not show this: WER is usually calculated after the text has been lowercased and stripped of punctuation. A model with a low error rate may handle commas poorly, so for subtitles this needs to be checked separately.
comments