Two kinds of voice, two kinds of licence
Library voices are synthetic voices the vendor offers to every customer. Check the licence scope: some cover online content but not broadcast or paid media, some exclude political or sensitive categories, and exclusivity is rarely available, so a competitor may use the same voice. Cloned voices are built from recordings of one person and need that person's consent covering purpose, media, territories, term and withdrawal, plus agreed payment; with professional talent, check whether union terms for digital replicas apply. Avoid sound-alikes. Imitating a famous voice in advertising lost in US courts long before AI, in Midler v. Ford (1988) and Waits v. Frito-Lay (1992).
Pronunciation is a brand asset
Synthetic voices stumble on exactly the words a brand cares about: its own name, product names, place names, acronyms that should be spelled out or said as words, prices and units. Look for a pronunciation dictionary that applies across every project, phonetic input using IPA, and support for the W3C's Speech Synthesis Markup Language, whose phoneme and substitution tags fix a word once. Keep lexicons per language, because a brand may choose to be pronounced differently in French and English markets. Add a pronunciation guide to the brand guidelines so every voice-over, human or synthetic, says the name the same way.
Languages, accents and delivery
A language on a vendor's list may have only a handful of voices, or voices that read cleanly but flatten emotion. Test the accents you actually need, such as Castilian or Latin American Spanish, European or Brazilian Portuguese, and Canadian or metropolitan French, and test long scripts, where some voices drift in pace or tone across twenty minutes of e-learning. Deliver at the sample rate and loudness each medium expects: 48 kHz for video, EBU R128 for European broadcast. Synthetic White produces voice-over and dubbing in 30 languages with timing preserved; native-speaker review before release remains good practice whichever tool you use.
Disclosure, phone calls and impersonation
Rules tighten as a synthetic voice gets closer to a real person or a live conversation. In the US, the FCC ruled in February 2024 that AI-generated voices in calls count as artificial voices under the Telephone Consumer Protection Act, so automated calls using them need prior consent. Tennessee's ELVIS Act, in force since July 2024, targets unauthorised AI replicas of a person's voice. In the EU, the AI Act has required deployers to disclose audio deep fakes since August 2026; providers must mark synthetic audio in a machine-readable way (by December 2026 for systems already on the market), and people talking to an AI system must be told so. UK radio ads still go through Radiocentre clearance.
Updated 25 September 2026