Recognition accuracy depends on your audio
Speech recognition quality is usually measured as word error rate, and it varies widely: by language, with widely spoken languages generally better served than smaller ones; by accent and dialect; and above all by the recording. Background music, crosstalk, echoing rooms and phone-quality audio all raise errors. Specialist vocabulary is the brand-specific weak spot: product names, drug names, model numbers and people's names. Look for custom vocabulary so brand terms are recognised, speaker labelling for interviews and panels, and sensible handling of numbers and currencies. Test with three clips: a clean presenter, a noisy event and your strongest accent.
Timing and readability rules
A transcript chopped into chunks at every pause is not a subtitle file. Subtitles have to stay on screen long enough to read: common style guides set adult reading speeds of about 160 to 180 words per minute, or up to 20 characters per second, and allow two lines of around 37 to 42 characters each. Lines should break at natural phrase boundaries, never between an article and its noun, and a subtitle should not hang over a shot change. Good generators apply these rules at export and let you set them per project, since children's content, e-learning and fast social edits call for different settings.
SRT, WebVTT and burned-in captions
SRT is the plain format: numbered cues with timestamps written as 00:00:01,000, with a comma before the milliseconds, and no standard styling. WebVTT starts with a WEBVTT header, uses a full stop in timestamps and supports positioning and styling, which is why browsers' built-in video players use it. Broadcast and streaming delivery often asks for TTML-based formats such as IMSC instead. Burned-in captions are part of the picture, so they always show, which suits muted autoplay on social feeds; sidecar files can be switched on and off and offered in several languages. Facebook, for instance, expects SRT files named with a language and country code, such as video.en_GB.srt.
Translation, accessibility and styling
Translated subtitles need re-segmenting, not line-by-line translation, because word order and length change between languages. Captions for deaf and hard-of-hearing viewers also carry speaker names and meaningful sounds, and WCAG makes captions for pre-recorded video a Level A requirement; the European Accessibility Act, applying since June 2025, brings WCAG-based rules to services such as e-commerce and banking. On social, styling is part of legibility: type size, a contrast box and a position clear of each app's interface. In Synthetic White, the video editor sits alongside dubbing into 30 languages with timing preserved and copy localised into 30 languages with terminology enforced.
Updated 25 September 2026