September 13, 2026
BUILD THE EVIDENCE LOOP: WHAT WILL REALITY TEACH YOU?

A live product teaches through user behavior, conversations, support requests, cost, retention, and results — not page views. Distribution (XIV), monetization (XV), and scale (XVI) converge here: without a loop that turns reality into decisions, shipping is theater. This lesson builds the loop.
The vocabulary of learning
- Hypothesis — a falsifiable statement about what must be true ("pilot analysts forward the brief without rewriting it").
- Signal — an observable trace of reality: a forward, a booking, a complaint, a bill.
- Metric — a counted signal over time: forward rate, completion rate, cost per brief.
- Qualitative evidence — what people say and why: interviews, support notes, review comments.
- Instrumentation — the logging and asking you set up before the test so the signal is captured.
- Feedback request — the direct question you ask one user at the right moment.
- Cohort — the named group you watch (e.g., "five pilot analysts, October").
- Decision — the recorded keep / change / kill choice the evidence forces.
- Experiment — the small test with one changed variable and a review date.
The compact loop runs six steps and changes one thing at a time:
state what must be true → ship a small test
→ collect behavior + comments → compare against hypothesis
→ record decision → change one thing
Changing one thing is the discipline. Two simultaneous changes produce two stories about any result, both plausible, neither provable. Small teams cannot afford unresolvable arguments; they can afford sequential tests.
The four-signal mix
Vanity metrics — views, impressions, signups — feel good and decide nothing. Every hypothesis in your loop needs four signals, one of each:
1. Behavior — what users did: completed the flow, forwarded the brief, booked, returned. 2. Quality / outcome — whether it was good: source-check pass rate, revision count, satisfaction with the after-state. 3. Direct question — what one customer said when asked: the interview quote, the support request, the reason for not returning. 4. Cost / effort — what it took: model spend per run, minutes of human review, support load.
A hypothesis passes only when behavior and quality agree and cost stays inside the guardrail; the direct question explains why. If behavior is strong but quality is weak (briefs forwarded, then corrected by clients), you have a trust problem, not a growth signal. If quality is strong but cost explodes ($4 per brief at pilot scale), you have a margin problem before you have a business.
For Sonariq's pilot: behavior (3 of 5 analysts forward without rewriting), quality (human source-check ≥ 95% claims linked), question ("what did you cut before forwarding?"), cost (≤ $1.20 plus ≤ 15 review minutes per brief). For the launcher: bookings completed unassisted, no-show rate, "what nearly stopped you?", cost per confirmed slot. Instrument before launch: log every run, timestamp every review, file every quote, tag every dollar.
Exercise: create LEARNING-LOOP.md
Create FINAL-PROJECT/LEARNING-LOOP.md with three hypotheses, each carrying the four-signal mix:
# LEARNING-LOOP — [Project], [cohort + dates]
## Hypothesis 1: [must-be-true statement]
- Behavior signal + metric + target: …
- Quality signal + target: …
- Direct question (asked when/whom): …
- Cost/effort signal + cap: …
- Collection method + instrumentation: …
- Review date: …
- If true → [decision + next test]; if false → [change-one + retest]; if mixed → [what decides]
## Hypothesis 2: …
## Hypothesis 3: …
## Decision log (append after each review)
- [date] Evidence seen → decision → one change → next review
Good first hypotheses test the riskiest beliefs, not the easiest wins: "the outcome matters enough to forward/pay/return," "quality survives without me in the room," "cost stays inside margin at 10× volume." Save channel and pricing optimization for hypotheses 4–6, after the core loop proves the job real.
Finish line: three hypotheses, each with four signals, collection method, review date, and pre-written true/false branches — a system that learns without pretending assumptions are facts.
Verify: for each hypothesis, ask "what instrument captures this signal today?" If the answer is "we will remember to check," add the log, the question script, or the dashboard line now. Then confirm the cohort is named and reachable — five specific people, not "users."
Common failure mode: the un-falsifiable hypothesis ("users will love the brief") with no target and no review date. Love is not a metric. Its mirror is signal-hoarding: twelve metrics, no decision. Four signals per hypothesis, one decision per review, one change after. The loop's output is decisions, not dashboards.
Pre-writing the decision
The most valuable line in LEARNING-LOOP.md is written before any data arrives: what each possible result would change. If forwards hit target but review minutes exceed the cap, the pre-written response might be simplify the draft template, not celebrate. If quality passes but nobody returns, the response is revisit the route from Lesson 90.8, not add features. Pre-commitment defeats the two classic evasions: moving the target after seeing the number, and collecting a fourth week of data instead of deciding. The review date is a promise to choose, and the decision log proves the loop runs rather than spins.
Check your understanding
1. Why must behavior, quality, question, and cost all be present before a keep/scale decision? 2. Rewrite "people want AI briefs" as a falsifiable hypothesis with a target and review date. 3. What breaks when two variables change between reviews?
Next
The loop tells you what reality thinks. Lesson 90.8 designs how reality finds you — one audience, one route, one fair exchange, built in from the start.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
