Guide

Where automatic captions break down

Automatic transcription is now good enough that the remaining mistakes are the interesting part. They are not scattered at random. They cluster in the same handful of places on almost every video, which makes them cheap to find once you know where to look.

Start free with Google 50 free credits. No credit card needed.

Names and anything invented

Speech recognition works by choosing the most probable sequence of words for a stretch of sound. A person's surname, a company, a product, or a piece of internal jargon is by definition improbable, so it gets replaced with something common that sounds close. This is the single most frequent error and the one your audience is most likely to notice, because those are the words they were listening for.

Two people at once

Overlapping speech is the hardest input there is. Interruptions, agreement noises, and the moment two people start a sentence together tend to produce a garbled line or a dropped one. Conversations, panels, and recorded meetings have far more of this than a scripted piece to camera, and it is worth scanning those sections specifically.

Numbers, units, and anything spelled out

Figures, dates, prices, percentages, and web addresses are transcribed as words rather than checked as values, so nothing catches an implausible one. A misheard digit reads perfectly well as a sentence, which is exactly what makes it dangerous in a business or technical video.

Audio that was never clean

  • A microphone across the room instead of near the mouth, which is the most common cause of a poor transcript by a wide margin.
  • Room echo, especially in a hard-surfaced office or a large hall.
  • Music or background noise sitting at a similar level to the speech.
  • Clipped or over-compressed audio, where the loud parts have already lost detail before anything is transcribed.
  • Phone speakers and laptop microphones, which roll off exactly the frequencies that distinguish similar consonants.

Improving the recording moves the result more than anything you can do afterwards. If a video is going to be captioned, that is a reason to get the microphone close on the day.

Accent, speed, and switching language

Recognition quality is uneven across accents, and it drops on very fast delivery and on trailing sentence endings. Switching between languages mid-sentence, which is normal in a lot of the world, is a particular weak point: the transcript is produced against one language and a borrowed phrase from another tends to come back as nonsense that looks like a typing error.

Why translation multiplies all of it

A translation is made from the transcript, not from the audio. An error in the transcript is therefore carried into the translated file as a confident, fluent, wrong sentence, and it is harder to spot there because you may not read the target language well enough to notice. Every job returns the original-language files alongside the translated ones, so read the two together: an error you can see in the language you speak is the same error sitting in the one you cannot. See how to translate subtitles well.

What to check, in order

  1. Every proper noun. Names, companies, products, and places, in that order of likelihood.
  2. The first fifteen seconds. It is the part almost everyone sees and most people judge.
  3. Numbers, dates, prices, and units.
  4. Any passage where more than one person is talking.
  5. The longest lines, which are usually long because something ran together that should not have.

Fixing what you find

The subtitle editor on Pro and Studio lets you correct the text and reburn the video. You can also edit the downloaded SRT in any text editor, since it is plain text; keep the encoding as UTF-8 or accented characters will break. A correction belongs to the job it was made in, because every job transcribes the audio from scratch. See what is an SRT file for the format.

Generate a transcript you can read through, with 50 free credits.

Continue with Google50 free credits. No credit card needed.

Frequently asked questions

Is a human transcript always better?

On difficult audio, yes, and it costs accordingly. On clear speech the gap is small enough that reading through a machine transcript and fixing the names gets you most of the way for a fraction of the effort.

Does a longer video transcribe worse?

Not inherently. Length matters because there is more to check, and because long recordings tend to be conversations, which contain more of the hard cases than a scripted piece does.

Why does it sometimes invent a plausible sentence?

Because it is choosing the most probable words for the sound it heard. When the audio is unclear it still returns its best guess, and that guess reads as fluent language rather than as an obvious gap.

Will fixing the text once fix it for other languages?

No. Each job transcribes the audio again, so a correction applies to the job where you made it. If wording matters, check the text on each run before you publish it.