September 13, 2026
IMAGE GENERATION IS ART DIRECTION, NOT MAGIC WORDS

You finished Part XII knowing how to turn a controlled build into a useful application. Part XIII asks a new question: can you make the visual asset that earns that application's first ten seconds of attention? A course lesson needs a hero image. A local launch needs a product photo that looks honest. A research story needs a visual metaphor that does not invent evidence. All three start in the same place — not a prompt box, but a brief.
The model renders; you direct
Bridge from Part IV: a model produces an output from inputs. It does not know your audience, your promise, or where the title text must sit. Your job is to define the outcome worth producing.
Think of image generation the way a film director thinks about a shot. The camera operator handles the tool. The director decides what the shot must communicate, what must be visible, what must never appear, and how it should feel. When the direction is vague — "make something premium" — every output is a surprise. When the direction is specific, outputs become comparable: does this one serve the goal better than that one?
That is why this class teaches a hierarchy, not a trick vocabulary:
communication goal → art direction → reference / material
→ generation instruction → selection / editing
Skip the first two steps and the last two cannot save you. A clever instruction on top of a fuzzy goal produces dramatic images that cannot be used. A clear goal with clear direction produces usable images even from a modest model.
The six terms to use correctly
- Creative brief: the job. Audience, message, promise, format, and constraints. Written before any tool opens.
- Audience: who must act, and what they care about. A ByeBuy student skimming lessons is not a shopper scrolling deals.
- Message: the one sentence the image must support. Not three sentences. One.
- Format: size, ratio, and placement. A desktop hero crop is a different job from a square feed card.
- Composition: where things sit in the frame — subject position, empty space for copy, foreground versus background.
- Subject / setting / visual style / reference / constraint / variation: the art-direction vocabulary. Subject is who or what is shown. Setting is where. Style is the visual grammar. Reference is a concrete example. Constraint is a hard rule. Variation is an intentional alternative to compare.
Learn these once and every later lesson in Part XIII — video shots, voice plans, UGC briefs — reuses them.
A real brief: the ByeBuy hero image
Here is the ByeBuy Classroom campaign hero brief we will reuse across this class:
- Goal: a student arriving at a lesson feels "this is a serious, practical course about building with AI" within three seconds.
- Audience: a builder skimming the course map, deciding whether to invest twenty minutes.
- Message: modern AI work is understandable and buildable.
- Format: wide 16:9 hero, safe on desktop, with generous negative space on the left third for the lesson title. No essential detail in the outer 10% (crop safety).
- Visual system: dark navy background, deep green secondary shapes, one restrained red accent. Clean, geometric, editorial — never neon, never cartoon, never photorealistic clutter.
- Subject: one clear focal object suggesting the lesson (for this lesson: an art-director's frame or viewfinder over a composed scene — not a robot holding a paintbrush).
- Constraints: no tiny text inside the image (all real text is added later in HTML), no brand logos, no recognizable living person, no cluttered background, readable at both desktop and phone widths.
Notice what the brief does: it answers *what must be true when the image succeeds* before anyone types an instruction. Anyone handed this brief could request the same image and get a recognizably similar result.
Why adjectives lose and structure wins
Beginners write prompts like this:
Every word is an adjective; none is a decision. *Premium* to whom? *Cinematic* in what ratio? What is the subject doing, where, with what light, with copy space where?
Compare a directed instruction built from the brief:
The second version names a subject (viewfinder over desk scene), an action (framing, implied), an environment (dark navy geometric background), a composition (subject right, copy space left), and constraints (no text, no logos, no people). The model finally has something to follow — and you have something to judge: is the left third actually empty? Is the red accent restrained? If not, you know exactly what to fix.
Rule of thumb: if you cannot sketch the image as three boxes on a napkin (subject box, background box, copy box), the direction is not ready.
The three running examples, directed
- Classroom campaign: the hero above. The test is legibility: title sits cleanly on the left at desktop and mobile, focal object survives the crop.
- Local launch: a bakery's sourdough launch. Subject: one loaf on a wooden counter, morning side-light, blurred shop interior behind, square crop, space at top for "Baked at 6 AM daily." Constraint: the actual loaf shape, crust color, and label must stay accurate — appetite appeal cannot change the product.
- Research story: an article about retrieval-augmented generation. Subject: a visual metaphor — a librarian's card drawer with one glowing card being pulled. Constraint: the chart/data stays in the article text; the image illustrates the idea of "finding the right card" and must not fake a screenshot, quote, or result.
Same hierarchy every time. Different goal, different constraints.
Exercise: write IMAGE-BRIEF.md
Pick one: a ByeBuy lesson hero, your local-launch product, or your research-story metaphor. Create IMAGE-BRIEF.md:
# IMAGE-BRIEF.md — [project name]
## 1. Goal & audience
- Viewer:
- One message the image supports:
- Action after seeing it:
## 2. Feeling & visual system
- Intended feeling (one line):
- Palette / materials:
- Style reference:
## 3. Format & placement
- Ratio & size: (16:9 hero / 1:1 feed / 9:16 short / portrait editorial)
- Where copy/title sits:
- Crop-safety notes:
## 4. Subject & scene
- Key subject:
- Action / arrangement:
- Setting / background:
- Camera / framing:
- Lighting:
## 5. Must-show / must-not-show
- Must show:
- Must NOT show: (tiny text, logos, real faces, invented claims...)
## 6. Selection test
- How I will pick the winner (3 checks):
Finish line: one IMAGE-BRIEF.md that another person could use to request the same image without calling you.
Verify fast: hand the brief to a friend or a fresh AI session and ask what image they would make. If they ask more than two clarifying questions, the brief is missing a decision — usually format, subject, or copy space. Common failure: writing the brief as adjectives ("eye-catching, vibrant") instead of placements ("subject right, empty left third").
Production gate (draft note — no art yet): treat any generation from this brief as a draft until human review passes facts, brand fit, legibility, accessibility (alt text, contrast, readable copy space), and permission (no real likeness, logo, or copyrighted reference without rights — use an original direction instead). Cover art is deferred to publish stage; keep this lesson's artifact as the brief text only.
Check your understanding
1. Recite the five-step hierarchy from memory. Where does a prompt box sit in it? 2. Why is "no tiny text in the image" a constraint rather than a style preference? 3. What three boxes would you sketch for the ByeBuy hero before generating anything? 4. Take a vague prompt you have used before. Which of subject, action, environment, composition, or constraints was missing?
Next
A brief says what is worth making. Next, Lesson 57.2 turns that brief into a controlled text-to-image direction — subject, camera, lighting, palette, ratio, exclusions — and teaches you to iterate one variable at a time.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
