September 13, 2026
VOICE GENERATION, CLONING, AND DUBBING: PERMISSION IS PART OF THE WORKFLOW

Lesson 61.1 taught you to write and mark a script worth performing. Now the harder question: whose voice performs it? A voice is not a font. It is personal, recognizable, and tied to a real human's identity and livelihood. Choosing a voice path is a creative decision and a consent decision at the same time.
By the end of this lesson, you can choose the right voice path for a project and record the choice in a VOICE-PLAN.md with source, language, owner, and review step.
Four paths: licensed synthetic, clone with permission, dub, human
Every voice project resolves to one of four options. Learn to name which one you are using:
| Path | What happens | Best when | Watch out for |
|---|---|---|---|
| Licensed synthetic voice | You select a ready-made AI voice under the tool's license terms | Anonymous explainers, prototypes, internal drafts, multilingual versions | License limits on commercial use; flat delivery on emotional material |
| Clone with permission | You replicate a specific real person's voice with their clear, written consent | A founder or host who wants a repeatable presence without re-recording | Scope creep — permission for one project is not permission forever |
| Dub / translate an existing recording | You re-voice recorded speech into another language or re-record it cleanly | International tutorials, localized lessons, accessibility tracks | Timing drift, mistranslated idioms, tone that no longer fits |
| Human performer | You hire or record a real person to read the script | High-trust, emotional, comedic, or brand-defining material | Scheduling and cost — but often the best result |
The technical differences matter less than beginners think. A good licensed voice with a strong script beats a sloppy clone of a famous-sounding voice every time. The differentiator is fit: does this voice match the message, the audience, and the consent you actually hold?
Why voice is personal: the no-endorsement-fake rule
Use a person's voice only with clear permission and project-appropriate terms. Write down three things before you generate a single line:
1. Who consented — the named individual, not "the team" or "the internet." 2. To what exactly — which script, which project, which languages, which distribution channels. 3. For how long — one video, one campaign, ongoing updates, and how consent can be withdrawn.
And one absolute boundary: never use a recognizable voice — cloned or imitated — to imply a real endorsement or a real conversation that did not happen. A synthetic founder "thanking" a customer who never spoke to them, a cloned celebrity "recommending" a product, a fake dialogue between two real people: these are deceptions, not productions. They destroy listener trust and can create serious legal exposure. If the listener would reasonably believe the person actually said it, you need that person's genuine statement or their explicit approval of the exact words.
When in doubt, pick the licensed synthetic voice and spend your energy on the script. Nobody is harmed by a clearly synthetic narrator reading honest words.
The multilingual workflow: meaning, timing, culture, reviewer
Dubbing is translation plus performance. A word-for-word translation read at the original timing usually fails — sentences run long, jokes land flat, names get mangled. Run this chain instead:
source script → translate for MEANING → adapt for TIMING and cultural fit
→ generate or record → FLUENT REVIEWER checks names, tone, and claims → publish
- Translate for meaning. Idioms, examples, and humor must be re-expressed, not transliterated. A baseball metaphor means nothing to many audiences; swap it.
- Adapt for timing. Some languages run 20–30% longer than English. Shorten sentences, allow scene re-timing, and never squeeze a rushed read under fixed visuals.
- Adapt for culture. Forms of address, politeness levels, currency, dates, and examples all shift. A "casual founder update" in one culture reads as disrespectful in another.
- Fluent reviewer, always. A fluent speaker checks the final audio — not just the translated text — for mispronounced names, wrong tone, and claims that changed meaning in translation. Machine translation pasted into a voice model is a draft, never a deliverable.
Budget the reviewer from the start. It is the cheapest quality step in the whole chain and the one most often skipped.
Three cases: which path and why
Practice the decision on three realistic briefs:
Case 1 — Anonymous product explainer. A 60-second "how it works" video for a new feature. No personal brand at stake, speed matters, budget is small. Choice: licensed synthetic voice. No consent complexity, easy to revise, easy to dub later. Spend the savings on the pronunciation sheet and the edit.
Case 2 — Founder narration. A two-minute story about why the company exists, told in first person as the founder. Trust is the entire point. Choice: clone with written permission — or the founder's real recorded voice. If the founder records once and approves a clone for future updates, scope the consent in writing: which channels, which topics, who approves each script. For the launch video itself, prefer the real recording; nuance is the message.
Case 3 — International tutorial. A 10-minute how-to that must ship in three languages. Accuracy and timing matter more than star power. Choice: human or licensed-voice source track, then dub with fluent review. Translate for meaning, re-time scenes, generate or record each language, and have a fluent reviewer sign off on names, claims, and tone before publishing.
Notice each answer names the permission and the quality logic together. That pairing is the skill.
Build it: the VOICE-PLAN.md
Your exercise: pick the voice path for all three cases above (or three of your own) and record them in one plan file:
VOICE-PLAN.md — [project name]
CASE 1 — [e.g. anonymous explainer]
- Voice source: [licensed synthetic / clone / dub / human + which voice]
- Audience + language: [who listens, in what language]
- Permission/approval owner: [named person who consented or approved]
- Pronunciation sheet: [5+ terms with say-as spellings]
- Review step: [who listens to the final audio and signs off]
CASE 2 — [e.g. founder narration]
(same five lines)
CASE 3 — [e.g. international tutorial]
(same five lines + fluent reviewer per language)
Worked miniature: for the tutorial case — source: founder's English recording; languages: English, Spanish, German; owner: founder (written consent for dubbing, dated); pronunciation sheet: five product terms per language; review: one fluent reviewer per language listens to final audio and checks names, numbers, and claims.
Check your understanding
1. Name the four voice paths and one situation where each is the best choice. 2. What three facts must be written down before cloning a real person's voice? 3. Why must a fluent reviewer check the final audio rather than just the translated text?
Common failure mode: a permission note that says "the founder is fine with it." Fine with what, where, for how long? Rewrite until a stranger could tell exactly what was consented to.
The finish line
You are done when you hold VOICE-PLAN.md: a voice source, audience and language, named approval owner, pronunciation sheet, and review step for each of three cases — with no endorsement fakes and no unreviewed dubs.
Next, Lesson 61.3 cleans up what the microphone caught: noise, room, edits, and the podcast pipeline.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
