September 13, 2026
90.E — EVIDENCE AND OPERATING REVIEW SCORECARD

User outcome, quality, cost, support, revenue, and next decision. Companions to Lessons 90.7 and 90.9. Run it weekly; append the decision log so learning compounds instead of evaporating.
The scorecard
# REVIEW-SCORECARD — [Project], week of [date], cohort: […]
## 1. User outcome (behavior)
- Signal/metric + target: … → observed: … → [pass / fail / mixed]
## 2. Quality
- Sample (rate + method): … → bar: … → observed: … → [pass/fail]
## 3. Direct voice (qualitative)
- Question asked (whom/when): … → sharpest quote: …
- Theme across quotes: …
## 4. Cost / effort
- Per-run: … (cap: …) / Period total: … (cap: …) / Review minutes: …
- Verdict: [inside / alert / pause-tripped]
## 5. Support
- Threads: … by category: … / Response promise met? [y/n]
- Top friction: …
## 6. Revenue / exchange (if testing)
- Offer test + conversions: … / Refund/cancel events: …
## Decision (one)
- Evidence → [keep / change-one / kill] → change: … → owner: … → next review: …
Scoring rules
Score behavior and quality together: both green means the job is real and good; behavior-green/quality-red means trust debt — pause scaling and fix the gate; both red means the job or route is wrong — change one, not five. Cost has veto power: any cap breach pauses the step that caused it before the next test. Qualitative evidence never votes alone but always explains the vote — file the quote that changed the decision verbatim.
Filled miniature (analyst pilot, week 2, illustrative)
Behavior: 3/5 forwarded unrewritten (target 3/5) → pass. Quality: 94% claims linked (bar 95%) → mixed. Voice: "I cut the peer table — numbers lacked dates." Cost: $1.10/brief + 18 review min (caps $1.20/15 min) → mixed. Support: 2 threads, both peer-table confusion. Decision: change-one — add date-stamped peer table template; owner Maya; re-review in 7 days.
From scorecard to backlog
Every mixed or failed row must produce exactly one backlog entry with owner and trigger date — never a vague resolution to do better. A quality miss becomes a gate-tightening task; a cost overrun becomes a cap or caching task; a support theme becomes a SPEC clarification. The scorecard week is only complete when the decision log points at these entries. Learning that never reaches TASKS.md is reporting, not operations.
Verify fast
Check instrumentation before scoring: every number traces to a log, a bill, or a filed quote — never memory. Confirm the cohort is still named and reachable. Then enforce the output rule: no review ends without one decision, one change, one owner, one date. Observations without decisions are minutes, not operations. Archive each scored sheet with the decision log so trends across weeks stay visible.
ARTICLE DISCUSSION
JOIN THE
CONVERSATION.
Got a question, a take, or a better way to do this? Log in and leave a comment.
