September 13, 2026
TEXT-TO-IMAGE: CONTROL SUBJECT, COMPOSITION, AND FORMAT

Lesson 57.1 gave you a brief: goal, audience, format, subject, constraints. Now you convert that brief into a generation instruction the model can actually follow — and you learn to iterate like a director instead of a slot-machine player.
New image versus edit: pick the right door
Text-to-image creates a new visual from a structured direction. Image editing (Class 58) changes an existing photograph — relighting a dim product shot, repairing damage, extending a background. Mixing them up causes real harm: generating a "better" product photo that quietly changes the label, or editing a chart into different data.
Rule: if the source photograph carries a factual claim — this is the product, this is the person, this is the data — start from the photo and edit conservatively. If no source photo exists and no claim depends on photographic truth — a hero illustration, a visual metaphor, a background plate — generate new. The local bakery's actual loaf? Photograph it. The ByeBuy hero's geometric viewfinder scene? Generate it. The research story's card-drawer metaphor? Generate it; keep the real chart in the article.
The nine parts of a working direction
Every strong direction answers the same nine questions, in this order:
1. Subject / action: who or what, doing what. One focal subject beats three competing ones. 2. Setting: where. Name the environment, not the mood of the environment. 3. Camera / framing: shot scale and viewpoint (wide, medium, close; eye-level, overhead; centered, rule-of-thirds right). 4. Lighting: source and quality (soft morning side-light, clean studio light, dim warm interior). 5. Palette / material: colors and surfaces (dark navy, brushed wood, matte paper — tie to your visual system). 6. Composition: placement of subject versus empty space. Always state where copy goes. 7. Style: the rendering grammar (flat vector editorial, clean product photo, soft 3D illustration). 8. Aspect ratio: the publishing decision (see below — set it at generation, not in post). 9. Exclusions: what must not appear (no text, no logos, no extra hands, no clutter).
Write them as short declarative lines, not a paragraph of adjectives. Models follow placed nouns better than stacked praise.
Ratio is a publishing decision, not a crop afterthought
Do not generate one square and stretch it everywhere. Each placement is a different composition:
| Placement | Ratio | What changes |
|---|---|---|
| Vertical short-form / Reel / TikTok | 9:16 | Subject large and centered; copy top or bottom clear of platform UI |
| Square feed card | 1:1 | Tight single subject; minimal background detail |
| Wide hero / video frame | 16:9 | Subject off-center; one full third reserved for title |
| Portrait editorial | 4:5 or 3:4 | Vertical subject (bottle, person, loaf); breathing room above headline |
Set the ratio in the tool before generating. A 16:9 hero re-cropped to 9:16 loses either the subject or the copy space — usually both. Your brief from 57.1 already names the ratio; honor it here.
Vague versus directed: the same bakery, twice
| Vague ("make a premium ad") | Directed (from the brief) |
|---|---|
| "Premium bakery ad, delicious, beautiful, 8k" | "Square 1:1 clean product photo. One round sourdough loaf centered on a light wooden counter, morning side-light from the left, blurred bakery interior behind. Top quarter soft empty blur for headline text. Warm natural palette, shallow depth of field. No text in image, no people, no extra pastries, no plastic packaging." |
| Result: random pastries, fake text, wrong loaf | Result: comparable, judgeable — is the loaf right? Is the top clear? Is the light morning? |
The directed version is longer because it makes decisions. Each decision becomes a review check later (Lesson 57.5). Copy this pattern for the Classroom hero and the research metaphor: subject, setting, camera, light, palette, composition, style, ratio, exclusions — every time.
Iterate one variable at a time
Amateurs change everything between attempts ("different style AND different angle AND add text"), then cannot say what improved. Directors test one variable:
1. Generate a baseline straight from the brief (2–3 outputs, same direction). 2. Pick the closest. Change one thing: composition (subject right instead of center), or metaphor (viewfinder versus stacked thumbnails), or lighting (studio versus morning). 3. Log each attempt: date, tool, direction text, ratio, seed/settings if available, what changed, score 1–5 on brief fit. 4. Stop at three deliberate variations. More than five and you are browsing, not directing.
Keep a reproduce log — a small Markdown table in the exploration folder. Without it, the winning direction evaporates and next week's "same style" request starts from zero. With it, Lesson 57.4's reusable prompt block writes itself.
Exercise: IMAGE-EXPLORATION/ with three outputs and a note
From your 57.1 brief, build this folder:
IMAGE-EXPLORATION/
brief.md (copy of IMAGE-BRIEF.md)
v1-baseline.png (+ v1-direction.txt + settings)
v2-composition.png (one variable changed)
v3-metaphor.png (a different variable changed)
SELECTION-NOTE.md
reproduce-log.md
SELECTION-NOTE.md template:
# Selection note — [project]
- Winner: v2 / v1 / v3
- Why it serves the audience (2 sentences):
- What fails in the rejected two (1 line each):
- Next step: (refine winner / reshoot direction / proceed to review in 57.5)
Finish line: an IMAGE-EXPLORATION/ folder with the brief, three outputs, a reproduce log, and a selection note naming a winner for audience reasons — not "it looks cool."
Verify fast: show all three to someone who has not seen the brief and ask which one matches the one-sentence message. If they pick wrong, the direction — not the model — needs work. Common failure: three variations that differ only in filter or seed. Force at least one composition change and one subject/metaphor change.
Check your understanding
1. When should you photograph or edit a source image instead of generating new? 2. Name the nine parts of a working direction without looking back. 3. Why must ratio be set before generation rather than fixed by cropping? 4. What does a reproduce log record, and which later lesson depends on it?
Next
You can now direct and iterate. Next, Lesson 57.3 widens the lens: the global field of image engines — American, Chinese, and open — and a durable ten-point framework for choosing the right one per job.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
