Sidecar files or burned-in text
Sidecar files sit next to the video and can be switched on, off or between languages. SRT is the simplest: numbered cues with timecodes that use a comma before the milliseconds. WebVTT (.vtt), the format browsers use with the HTML track element, adds styling and positioning and uses a full stop in timecodes. Broadcasters and streaming services often want TTML-based formats instead. Burned-in, or open, captions are part of the picture: they show in every feed and player, which suits muted social video, but they cannot be turned off, resized or swapped for another language, and every language needs its own render.
Reading speed and line length
Captions are read, not heard, so they must keep pace with readers. A common ceiling is about 20 characters per second for adult viewers and 17 for children's content, with lower limits for audiences or languages that read more slowly. Keep to two lines, around 42 characters per line for Latin scripts and far fewer for Japanese or Chinese, and break lines at natural phrase boundaries, never between an article and its noun. Leave each caption up long enough to read, avoid running across shot changes, and condense fast speech rather than rushing the reader. Verbatim is not always better.
Accessibility duties
For deaf and hard-of-hearing viewers, captions carry more than dialogue: speaker identification, and sound cues such as [door slams] or [music stops] when they matter. WCAG requires captions for prerecorded video at level A. The European Accessibility Act (Directive (EU) 2019/882) has applied since 28 June 2025 to services including e-commerce, consumer banking and access to audiovisual media, with EN 301 549 the usual technical benchmark; micro-enterprises providing services are exempt. For an in-scope service such as an online shop, video on its own website or app is part of what must be accessible, so caption it. In the US, programmes shown on television with captions must also be captioned when put online, under the CVAA and FCC rules.
Translating captions
Translate from the timed, corrected source file, not from a fresh transcript, so timing and line breaks start from something checked. Then re-time: translations usually run longer, and a German caption that exceeds the reading speed needs condensing, not a longer display that collides with the next line. Right-to-left languages need files and players that handle direction correctly. Speech recognition stumbles on brand names, product codes, accents and jargon, so correct the source before anything is translated. Synthetic White localises copy into 30 languages with terminology and approved claims enforced, so product names and claim wording follow the approved terms in each language.
Updated 25 September 2026