Clean-up and editing without changing meaning
Automated clean-up now handles noise, room echo, sibilance, level differences between speakers and the long pauses of remote recording. Editing is where judgement enters. Removing filler words and false starts makes a guest sound sharper; removing a hesitation before a difficult answer can change what the answer means. Leave edits of substance to a human producer, and send guests anything contentious before release. Finish to a consistent loudness: Apple's guidance for podcasts is around −16 LUFS with peaks no higher than −1 dBFS, and matching that across episodes stops listeners reaching for the volume.
Clips, show notes and chapters
Clips work when they stand alone: a complete thought, usually under a minute, with captions burned in and the speaker named on screen, cut in 9:16 for short-form platforms and 1:1 for feeds. Never clip a guest in a way that reverses their point. Show notes should summarise, credit guests and link what was mentioned. Chapters help listeners navigate: MP3 files can carry them in ID3 chapter frames, the Podcasting 2.0 namespace adds a chapters tag to the RSS feed, and YouTube builds chapters from description timestamps that start at 0:00, with at least three chapters of ten seconds or more.
Transcripts worth publishing
A transcript makes an episode accessible to deaf and hard-of-hearing listeners, searchable on the show's website and quotable without re-listening. Speech recognition handles ordinary words well and the important ones badly: guest names, company names, figures and jargon. Correct those before publishing, mark speakers, and keep the corrected file as the source for show notes, clips and translations. The Podcasting 2.0 transcript tag lets apps that support it display a creator-supplied transcript, in formats such as SRT or VTT. For interviews that touch regulated claims, the transcript is also the record of what was actually said.
Voice cloning, consent and disclosure
Cloning a host's voice can fix a misspoken word, voice an ad read in another language or keep a show going through an illness. It needs written consent that names each of those uses, sets a term, gives the host approval over scripts and ends with deletion of the voice model. Tennessee's ELVIS Act of 2024 made unauthorised use of a person's voice, including AI imitations, actionable, and other states and countries protect likeness in similar ways. Tell listeners when a segment uses a synthetic voice. Host-read ads carry extra risk: in the US, the FTC expects endorsements to reflect the endorser's honest views, so a cloned read the host never approved is a problem.
Updated 25 September 2026