Formats and the companion
Music streaming ads are usually 15 or 30-second spots between tracks, often with a companion image on screen that listeners can tap while the ad plays. Podcast ads come as pre-roll, mid-roll and post-roll, either baked into the episode or inserted dynamically at download. Delivery runs on the IAB's VAST standard, which absorbed the audio-specific DAAST in version 4.1, and the IAB Tech Lab's podcast measurement guidelines define what counts as a download. Writing for audio is its own craft: one message, the brand named early and repeated, a call to action people can remember without a screen, such as a short web address or a code.
Host-read or produced
Host-read ads borrow the host's credibility and voice; produced spots are consistent and easy to scale. AI fits the produced side well: scripts per segment, synthetic or licensed voices in many languages, music beds and fast versioning. It should not fake a host read. Generating a real host's voice without explicit consent is a likeness problem, and in the US it can breach state law such as Tennessee's ELVIS Act of 2024. Paid reads must also be recognisable as advertising: the FTC's Endorsement Guides, revised in 2023, apply in the US, and in the UK the CAP Code requires marketing to be obviously identifiable.
Loudness and the mix
Streaming services and podcast apps normalise loudness, turning loud masters down to a common level. A spot squashed to sound loud therefore ends up no louder than its neighbours, only flatter and more tiring. For podcasts, Apple recommends about -16 LKFS with a tolerance of 1 dB and true peaks no higher than -1 dBFS, and ad platforms publish targets of their own. Mix for earbuds and car speakers, and check every spot on a phone speaker. Synthetic voices need direction as much as human ones: pace, emphasis and breaths decide whether thirty seconds sound like a person talking or a machine reading.
Voice consent and the records behind it
A voice is part of a person's identity, and synthetic use needs a contract that says so: which languages, which media, which markets and for how long. Stock voices need licence terms that cover advertising. From August 2026 the EU AI Act requires disclosure of realistic synthetic audio that could be taken for a real person. The studio voices and dubs spots into 30 languages with timing preserved, uses licensed music beds, and keeps a per-asset record of which model produced each voice, so questions about consent, disclosure and licensing have documented answers rather than recollections.
Updated 25 September 2026