Use case · AI content creation

AI subtitles and captions that meet the standard.

AI can transcribe, time and translate subtitles in minutes, but good captions follow rules that speech recognition does not know: a reading speed of roughly 17 to 20 characters per second depending on audience and language, short lines broken at natural points, and speaker and sound cues for deaf viewers. Choose sidecar files or burned-in text per platform, and check every name and claim.

Sidecar files or burned-in text

Sidecar files sit next to the video and can be switched on, off or between languages. SRT is the simplest: numbered cues with timecodes that use a comma before the milliseconds. WebVTT (.vtt), the format browsers use with the HTML track element, adds styling and positioning and uses a full stop in timecodes. Broadcasters and streaming services often want TTML-based formats instead. Burned-in, or open, captions are part of the picture: they show in every feed and player, which suits muted social video, but they cannot be turned off, resized or swapped for another language, and every language needs its own render.

Reading speed and line length

Captions are read, not heard, so they must keep pace with readers. A common ceiling is about 20 characters per second for adult viewers and 17 for children's content, with lower limits for audiences or languages that read more slowly. Keep to two lines, around 42 characters per line for Latin scripts and far fewer for Japanese or Chinese, and break lines at natural phrase boundaries, never between an article and its noun. Leave each caption up long enough to read, avoid running across shot changes, and condense fast speech rather than rushing the reader. Verbatim is not always better.

Accessibility duties

For deaf and hard-of-hearing viewers, captions carry more than dialogue: speaker identification, and sound cues such as [door slams] or [music stops] when they matter. WCAG requires captions for prerecorded video at level A. The European Accessibility Act (Directive (EU) 2019/882) has applied since 28 June 2025 to services including e-commerce, consumer banking and access to audiovisual media, with EN 301 549 the usual technical benchmark; micro-enterprises providing services are exempt. For an in-scope service such as an online shop, video on its own website or app is part of what must be accessible, so caption it. In the US, programmes shown on television with captions must also be captioned when put online, under the CVAA and FCC rules.

Translating captions

Translate from the timed, corrected source file, not from a fresh transcript, so timing and line breaks start from something checked. Then re-time: translations usually run longer, and a German caption that exceeds the reading speed needs condensing, not a longer display that collides with the next line. Right-to-left languages need files and players that handle direction correctly. Speech recognition stumbles on brand names, product codes, accents and jargon, so correct the source before anything is translated. Synthetic White localises copy into 30 languages with terminology and approved claims enforced, so product names and claim wording follow the approved terms in each language.

Updated 25 September 2026

Questions

AI subtitles and captions, answered.

What is the difference between subtitles and captions?

Subtitles translate or transcribe dialogue for viewers who can hear the audio. Captions, often called subtitles for the deaf and hard of hearing, also identify speakers and describe relevant sounds and music. Accessibility standards ask for captions, so marketing video meant for everyone should be captioned, not just subtitled.

Should we use SRT or VTT files?

Use what the platform asks for. SRT is the most widely accepted and easiest to edit; WebVTT is the web standard for HTML5 players and supports positioning and styling. Most tools convert between them, but check the timecode separator and the character encoding, which are the usual causes of failed uploads.

How accurate are AI-generated captions?

Accurate enough to start from, not to publish unchecked. Speech recognition does well on clear speech and common words and worst on brand names, technical terms, strong accents, overlapping speakers and noisy audio. For regulated claims, prices or medical language, a human check of every caption is part of compliance, not polish.

Get started

See your AI studio generating this week.

Thirty minutes. A live studio in your colours, your brand hub loaded, a real campaign brief.

See plans