September 13, 2026
AUDIO CLEANUP, SOUND DESIGN, AND PODCAST PRODUCTION

Lessons 61.1 and 61.2 gave you a script and a consented voice. Now reality intrudes: the recording has air-conditioner hum, the room echoes, one guest is quiet, the dog barked at minute four. Cleanup and sound design decide whether anyone finishes listening.
By the end of this lesson, you can take a raw 2-minute recording and produce an edit decision list plus a clean, reviewed short audio export.
The vocabulary: noise, room, EQ, compression, loudness, bed, SFX, mix
Eight terms cover most podcast and video-audio conversations:
| Term | What it means | What it fixes or adds |
|---|---|---|
| Noise reduction | Removing steady background hiss, hum, or fan noise | A kitchen-table interview that sounds studio-adjacent |
| Room tone | The natural ambience of a space; echo and reverb character | Matching a pickup line so it does not sound pasted in |
| EQ (equalization) | Boosting or cutting frequency ranges | Removing muddiness, taming harsh "s" sounds |
| Compression | Narrowing the gap between loud and quiet moments | A guest who whispers then laughs without deafening anyone |
| Loudness | Overall perceived volume, standardized per platform | An episode that is not suddenly quiet after the ad |
| Music bed | Low background music under speech | Energy and continuity under an intro or montage |
| Sound effect (SFX) | A short discrete sound: click, whoosh, chime | Signaling a transition or illustrating an action |
| Mix | The final balance of voice, music, and effects | Everything audible, nothing fighting the narration |
You do not need to master each knob. You need to hear what each one does and know which problem each solves — then let tools or a collaborator execute.
What AI cleanup does well — and where it damages
AI cleanup tools are genuinely strong at five jobs: removing steady background noise, transcribing speech to text, suggesting edit points from the transcript, isolating speech from mixed audio, and generating draft assemblies. A noisy café interview can become usable in minutes.
But every cleanup pass can damage the thing it saves. Listen for artifacts after each pass:
- Watery or robotic tails on words — noise reduction pushed too far.
- Clipped word starts — consonants eaten by aggressive gating.
- Pumping music — the bed ducking unnaturally under speech.
- Flat, lifeless voices — over-compression removed all dynamics.
The rule: clean in small steps and A/B every change against the original. If a listener can hear the processing, you went too far. A little honest room tone beats a sterile, artifact-ridden voice every time.
The pipeline: record to clips
Run every spoken-audio project through the same seven-stage chain:
record → transcribe → edit STORY → clean AUDIO → add SOUND → review → export clips + full episode
Note the order: story before sound. Cut the content first — remove tangents, reorder for clarity, tighten the argument — using the transcript as your editing surface. Only then clean the audio and add music and effects. Beginners do it backwards: they polish the EQ on a paragraph that should have been deleted.
Stage by stage:
1. Record — capture cleanly: close microphones, quiet room, separate tracks per speaker when possible. 2. Transcribe — generate a full transcript with speaker labels and timestamps. 3. Edit story — cut, reorder, and tighten from the transcript. Decide what the listener learns and in what order. 4. Clean audio — noise reduction, EQ, compression, loudness, in that order, checking artifacts each time. 5. Add sound — licensed or original beds and effects only where they support comprehension (transitions, emphasis, energy). 6. Review — a full listen on headphones and one cheap speaker; verify every name, number, and claim. 7. Export — the full episode plus short clips cut for sharing, each with captions.
The transcript is not the script
A transcript records what was said, including the ums, the false starts, the wrong numbers corrected mid-sentence, and the guest's confident misstatement of a date. A script states what you stand behind. Converting one into the other requires verification:
- Names — spell every person, product, and place name; check against sources.
- Numbers — every statistic, price, date, and percentage re-checked. Speech mishears numbers constantly.
- Claims — any factual assertion ("studies show," "our users save...") traced to its source or cut.
- Captions — generated captions inherit every transcript error and add timing errors. Proofread them as a separate pass.
If a claim cannot be verified, cut it or downgrade it to what you know: "one customer told us" instead of "customers save 40%."
Build it: the 2-minute edit plan
Your exercise: take any raw 2-minute recording — a voice memo, a meeting excerpt, a test narration — and write an edit decision list:
EDIT PLAN — [recording name] — raw duration [2:00]
REMOVE ([timestamps]):
- [0:12–0:25] tangent about scheduling — breaks the argument
KEEP ([timestamps]):
- [0:26–1:10] core explanation — the one idea the clip exists for
CLARIFY (re-record or pickup):
- [1:10] number stated as "forty" — verify; sounds like "fourteen"
CAPTIONS:
- [proofread all lines; fix 3 names/terms; check timing at 0:45 overlap]
SOUND (only where it supports comprehension):
- [soft bed under 0:00–0:08 intro, fade before speech; one transition sting at 1:10]
REVIEW: [listener name] checks names/numbers/claims; artifact listen on headphones
EXPORT: [full 2-min clean WAV/MP3 + one 0:30 clip with captions]
Worked miniature: a founder's rambling 2-minute feature description becomes a 75-second keeper. Remove the 20-second apology for the slides. Keep the single workflow demo. Clarify the pricing line with a pickup recorded in the same room. Add one soft bed under the intro only. Result: a clip short enough to share and clean enough to trust.
Check your understanding
1. Why does story editing come before audio cleaning in the pipeline? 2. Name two cleanup artifacts and what over-processing step causes each. 3. A guest states a statistic confidently on tape. What do you do before publishing it?
Common failure mode: the over-polished ramble — pristine EQ and a lovely bed under four minutes that say one minute's worth of things. If your remove-list is empty, you edited sound, not story. Cut first.
The finish line
You are done when you hold an edit decision list (remove, keep, clarify, caption, sound) and a clean, reviewed short audio export with verified names, numbers, and captions.
Next, Lesson 61.4 scores the picture: music and sound effects chosen for fit and rights, not just mood.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
