September 13, 2026
VOICE IS AN INTERFACE: SCRIPT, PERFORMANCE, AND LISTENER TRUST

Class 60 showed you how moving pictures earn belief: previs before spending, proof before claims. Sound works the same way, only faster. A listener decides within seconds whether a voice is clear, honest, and worth following — long before they judge the visuals. This class teaches you to direct that voice, whether it comes from a human throat or a synthetic model.
By the end of this lesson, you can turn 250 written words into a marked-up 45–60 second narration script ready for a human or synthetic voice read.
The vocabulary: TTS, clone, narration, dubbing, and direction
Get the terms straight first, because the rest of the class builds on them:
| Term | What it means | Example |
|---|---|---|
| Text-to-speech (TTS) | Software renders written text into spoken audio using a synthetic voice | A course lesson narrated by a licensed synthetic voice |
| Voice clone | A model trained or tuned to sound like a specific real person | A founder's approved voice reading weekly updates |
| Narration | A voice that explains, guides, or tells a story over visuals | A ByeBuy lesson narrator walking through a concept |
| Dubbing | Replacing or adding a voice track in another language (or re-recording the same one) | A tutorial re-voiced from English into Spanish |
| Performance direction | The human decisions about pacing, emphasis, tone, and energy | "Slower here. Land the number. Smile on the last line." |
| Pacing | The speed and rhythm of delivery — sentence length, pauses, acceleration | Short sentences for urgency; long ones for explanation |
| Pronunciation guide | A written list of exactly how names, terms, and numbers should sound | "ByeBuy (bye-BUY), Supabase (SOO-pah-bass)" |
| Pickup | A short re-record of one line or section to fix an error | Re-reading only sentence three, not the whole script |
A synthetic voice is an instrument. Direction is the playing. Beginners obsess over which voice model to pick; professionals obsess over the script and the markup, because those decide the result no matter which voice reads it.
Script first: writing for the ear is not writing for the page
A polished voice cannot rescue a vague script. In fact, a beautiful synthetic read of a muddled paragraph sounds worse than an honest human stumble — the smoothness signals confidence the words have not earned.
Writing for the page rewards density: long sentences, nested clauses, terms the reader can re-read. Writing for the ear rewards the opposite:
- One idea per sentence. Listeners cannot scroll back. If a sentence has two claims, split it.
- Short sentences, varied rhythm. Three medium sentences followed by one short punch line keeps attention.
- Concrete words first, terms second. "The tool that remembers what you wrote" lands before "persistent vector memory."
- Spoken numbers and names. Write "twenty twenty-six" or "two thousand twenty-six" exactly as it should sound, not "2026" and hope.
- Signposts. "Here is why that matters." "Three things happen next." Ears need road signs; eyes have paragraphs.
Read every draft aloud before you record or generate anything. If you stumble, the listener will too. Fix the sentence, not the voice.
Markup: pause, emphasis, pronunciation, length, tone
A narration script is a score, not an essay. Mark it so any reader — human or machine — performs it the same way. Use a simple, visible markup:
- [pause] or [beat] — a short silence. Use it before a key claim and after one, so the point lands.
- **CAPS or \*asterisks\*** — emphasis. Mark the one word per sentence that carries the meaning, never five.
- [say: ...] — pronunciation. Spell out exactly how a tricky term sounds:
[say: "SOO-pah-bass"]. - [slow] / [quick] — pacing notes for a section. Slow down for definitions and numbers; quicken for lists and energy.
- [tone: ...] — emotional direction: warm, neutral, serious, curious. One tone per section, changed deliberately.
Most TTS tools also accept their own syntax (breaks, emphasis tags), but keep your master script in this human-readable form first. Tool syntax changes; a marked script ports anywhere.
The Classroom narrator pattern
ByeBuy lessons use a repeatable narrator shape. Steal it for your own explainers:
1. Concise opening — one or two sentences that name the topic and the payoff. 2. Concept explanation — the idea in plain language, with the real term introduced once. 3. Example — one concrete story that proves the concept works. 4. Recap — one or two sentences restating what the listener now knows. 5. Cue to the screen — direction to what the viewer should look at next: "Watch the diagram as the three steps light up."
The voice never carries the whole lesson alone. It opens, explains, proves, recaps, and hands off to the visual. Lessons 61.3 through 61.5 will add cleanup, music, and presenters to this spine — but the spine is the script.
Build it: 250 words to a 45–60 second script
Your exercise: take roughly 250 written words of explanation — a paragraph from your own project, a product description, a lesson draft — and convert it into a narration script.
Rules of thumb: English narration runs about 140–160 words per minute, so 45–60 seconds means roughly 110–150 spoken words. You will cut. That is the point.
Use this template for your NARRATION-SCRIPT.md:
NARRATION-SCRIPT.md — [title] — ~[word count] words, ~[seconds]s target
VOICE: [human name or synthetic voice + backup choice]
TONE: [one line, e.g. warm, steady, curious]
PRONUNCIATION GUIDE (5 terms):
1. [term] — [say: "..."]
2. ...
3. ...
4. ...
5. ...
SCRIPT (mark 3+ pauses with [pause], emphasis with *stars*):
[Opening — 1-2 sentences]
[pause]
[Concept — plain language]
[Example — concrete]
[Recap + cue — what to look at next]
Worked miniature: a 250-word draft about database indexes becomes a 130-word script. Opening: "Searching a thousand rows is instant. Searching ten million is not — *unless* the database built an index." [pause] Concept: "An index is a *sorted shortcut* [say: ...] the database maintains beside the table." Example: "ByeBuy's product search dropped from four seconds to a blink after one index on the name column." Recap plus cue: "So: an index trades a little storage for a lot of speed. [pause] Watch the diagram as the query skips the full scan."
Check your understanding
1. Why does a polished voice make a vague script sound worse, not better? 2. Name the five markup types and say what each one controls. 3. Your script runs 200 words. Roughly how many seconds is that, and what should you cut first?
Common failure mode: marking everything — five emphasized words per sentence, pauses everywhere, three tone changes in 40 seconds. If every line is special, nothing is. Mark less, then listen: does the one key claim in each section land?
The finish line
You are done when you hold NARRATION-SCRIPT.md: a 45–60 second script with three marked pauses, five pronunciation entries, emphasis and tone notes, and a cue to the screen — ready for a human or synthetic voice to read cold.
Next, Lesson 61.2 tackles the permission question: whose voice is this, and what gives you the right to use it?
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
