Use case
Subtitles for recorded talks and webinars
A talk is recorded once and then usually sits somewhere unwatched. Captions and a transcript are what make it findable, usable by people who cannot hear it, and worth anything to an audience that does not share the speaker's language.
Why a recording needs them
- Accessibility. A recorded talk published without captions is unusable for part of its audience, and for public bodies and universities that is usually an obligation rather than a preference.
- Technical vocabulary. Specialist terms are hard to catch by ear, particularly in a second language, and seeing them written is often the difference between following a talk and losing it.
- Room audio. Conference recordings are made in hard rooms with distant microphones, so the sound is exactly the kind that people struggle with.
- Searchability. A transcript is text, and text is the only part of a video that can be searched or quoted.
How to do it
Upload the recording
MP4, MOV, AVI, or MKV up to 500 MB, or paste a link from a supported source.
Choose the language
The spoken language for captions and a transcript, or any of the 30 target languages for a version aimed at another audience.
Read the transcript before publishing
Speaker names, institutions, and technical terms are where an automatic transcript is weakest, and a talk is full of all three.
Publish the video and the text
You get the subtitled video, an SRT with timings, and a plain TXT transcript in both languages.
Terminology is worth a pass
Every field has words that are common inside it and vanishingly rare outside, which is the precise condition under which speech recognition substitutes something more probable. A talk can be otherwise clean and still get its central term wrong throughout. Fix it once in the editor on Pro and Studio, or in the downloaded SRT, and see why automatic captions get things wrong for the rest of the pattern.
Reaching people who were not there
The audience for a recorded talk is larger than the room and less likely to share its language. Translated subtitles are what turn a local event into something with an international audience, and the marginal cost of another language is one more job on the same file. See all 30 languages.
Clips travel further than the full recording
A forty-minute talk gets shared far less than the four minutes that made the point. Trim controls on Pro and Studio let you cut a segment before subtitling it, which also keeps you inside the 500 MB upload limit and spends fewer credits. Publish the full recording for the people who want it and clips for everyone else.
Captions, not just subtitles
If the reason for doing this is accessibility, the distinction matters: captions are written for someone who cannot hear the audio and carry more than dialogue. Automatic subtitles are a strong starting point rather than a finished caption track. Subtitles vs captions sets out the difference, and accessible video captions covers what a usable track needs.
Caption a recorded talk with 50 free credits.
Continue with Google50 free credits. No credit card needed.Frequently asked questions
Can I subtitle a full conference recording?
Uploads are capped at 500 MB per file, so a full session usually needs cutting down first. That tends to produce something more watchable anyway.
Will it identify who is speaking?
No. It transcribes speech without labelling speakers, which matters most for panels and question sessions. Labels can be added to the downloaded SRT, which is plain text.
How well does it handle audience questions?
That is the weakest audio in any recording, since questions are usually asked away from a microphone. Expect to correct those sections by hand if they matter.
Is this enough for an accessibility requirement?
It gives you an accurate starting point quickly, but a compliant caption track needs checking and usually needs speaker and sound information adding. Read the transcript rather than publishing it unseen.