ByeBuy.ai
BUILD YOUR ESCAPE ROUTE · ✦ CURSOR · HOST IT · ◫ SUPABASE · CONNECT IT · ↯ RELAY · BUILD YOUR ESCAPE ROUTE · ✦ CURSOR · HOST IT · ◫ SUPABASE · CONNECT IT · ↯ RELAY ·
CURRICULUM
← BYEBUY NOTES

September 13, 2026

HUMAN-IN-THE-LOOP IS A DESIGN CHOICE

ByeBuy.ai artwork for Human-in-the-Loop Is a Design Choice
ByeBuy.ai lesson artwork

Lesson 84.3 ended every contract at the same place: exceptions go to a human. That sentence is easy to write and easy to make meaningless — a queue nobody checks, an "approval" that gets rubber-stamped, an escalation to an inbox nobody owns. This lesson designs the human side deliberately: where people decide, what they see when they decide, and how much freedom the machine has before it must stop and ask.

Human-in-the-loop is not a feeling. It is a set of mechanisms.

The mechanisms

  • A review queue is where prepared work waits for a decision. It has an owner, a service target (decide within N hours), and a visible age. A queue with no target is a oubliette.
  • An approval gate is a step nothing passes without a recorded human decision: publish, refund, promise.
  • Confidence is the workflow's self-reported certainty — and you should treat it as a routing hint, never as truth. Low confidence routes to review; high confidence still gets sampled.
  • An exception is anything the contract did not anticipate. An escalation is the exception's path to a person with a deadline.
  • An audit is a periodic look back at runs — including the confident, uneventful ones — to catch drift before customers do.

The five-level ladder

Place each workflow step on exactly one rung:

1. Manual. A person does the work; AI may answer questions. 2. AI-assisted. The person works; AI drafts, summarizes, or checks inside the person's view. 3. AI-prepared, human-approved. AI produces the artifact; nothing moves without a recorded approval. 4. Bounded autonomous. The workflow acts alone inside the contract's bounds; exceptions escalate; a sample is audited. 5. Audited autonomy. Routine runs need no per-run approval, but full logs exist and audits happen on schedule.

Movement up the ladder is earned, never assumed. A step climbs only after quality is visible at its current rung for a full review cycle: low exception rates, clean audits, no surprises. New steps start at 1–3. Anything touching the outside world — customers, money, public claims — starts no higher than 3.

Teams fail here in two directions. The fearful keep everything manual and drown. The hurried jump to 4 or 5 because the demo worked twice. Both skip the evidence step. The ladder exists to force the question: what did the last thirty runs prove?

What always gets review

Six categories stay behind a human gate regardless of confidence scores:

  • Public claims made in your name.
  • Promises to customers (dates, prices, guarantees).
  • Money movement in either direction.
  • Account changes: access, permissions, ownership.
  • Ambiguous source material where evidence conflicts.
  • Unusual cases the contract never anticipated — the first-of-a-kind run.

The logic is uniform: the cost of an error dwarfs the cost of a review. A wrong refund is reversible with embarrassment; a wrong public claim or account change can be neither. When in doubt, the step is amber until a full cycle proves otherwise.

Design the queue itself with the same care as the workflow. A good review queue shows the reviewer exactly what changed and why: the input, the prepared output, a diff against the previous version, the sources cited, and the workflow's stated confidence with reasons. It batches sensibly — urgent exceptions interrupt, routine approvals collect into one or two daily sessions — and it makes the decision one click plus an optional note, recorded in the log. A queue that shows only conclusions ("approve this refund?") trains reviewers to click without reading. Show the evidence and you get judgment; show the conclusion and you get a rubber stamp with extra steps.

Exercise: mark every step green, amber, red

Take your operation map (84.1) and contract (84.3). Mark each step:

  • Green — safe to automate at rung 4–5: checkable, reversible, low blast radius.
  • Amber — prepare but require review at rung 3: AI drafts, human approves.
  • Red — human owner decides at rung 1–2: the six categories above.

Write it as a short table with one-line reasons:

| Step | Mark | Rung | Reason |

|------|------|------|--------|

| Validate source completeness | Green | 4 | Checkable against schema; reversible |

| Draft customer reply | Amber | 3 | Promise risk; needs approval |

| Issue refund | Red | 2 | Money movement; owner decides |

Finish line: a deliberate human boundary — every step marked, every amber and red step naming its reviewer and decision target.

Verify: confirm no red-category item sits at green, and every amber step has a queue with an owner and a service target. An amber step with nowhere to wait is a green step wearing a costume.

Common failure mode: rubber-stamping — approving a hundred AI-prepared items without reading. Counter it with sampling audits and by making the reviewer see the evidence (source, diff, confidence reason), not just the conclusion.

Check your understanding

1. What evidence must a step show before climbing from rung 3 to rung 4? 2. Why is confidence a routing hint rather than a decision-maker? 3. Name the six categories that always get human review, and the reasoning they share.

Next, Lesson 84.5: boundaries set, build the minimum viable operations dashboard — five metrics and a weekly question for each, so the system tells you when it needs attention.

ARTICLE DISCUSSION

JOIN THE
CONVERSATION.

0 COMMENTS

BYEBUY ACCOUNT ACCESS

Sign in

Use your account to save routes and make the catalogue yours.

Enter your email and we’ll send a secure sign-in link and code.

NEW ROUTES ADDED WEEKLY · 9,235 CATALOGUE ENTRIES · BUILD · DEPLOY · QUERY · STACK · SAY BYE TO BUY · NEW ROUTES ADDED WEEKLY · 9,235 CATALOGUE ENTRIES · BUILD · DEPLOY · QUERY · STACK · SAY BYE TO BUY ·