FINAL — YOUR BOX SCORE
HOW TO READ THIS: green = you and the booth agreed · amber = one grade apart, a judgment call · red = a blown call worth a second look. In a real support org, the red and amber plays are exactly what a human QA lead reviews next — the booth surfaces them, a person decides. AI as the replay official, never the umpire.
HOW THE CALLS WORK
Four calls, no sliders
The old version of this page asked you to grade five rubric categories on every transcript. Nobody QAs like that for fun. Now it's one call per play — Web Gem (flawless), Base Hit (does the job), Error (fixable miss), Ejection (policy breach) — the same four-point scale a real rubric collapses to, wearing a much better uniform.
The replay booth
Every play carries the booth's own call and its reasoning, written in advance the way an AI QA model annotates real tickets. The reveal-after-you-commit order matters: you're never anchored by the machine. That's how AI-assisted QA should work in production too.
Watch for landmines
A few plays hide an automatic error — leaked data, a promise the company can't keep, a policy breach dressed up in a friendly tone. A response can read warm and still be an Ejection. The booth will let you hear about it.
All transcripts are fictional and were written with Claude for this game — no real customers, companies, or tickets. Unofficial fan project of good support everywhere. Calibration is a skill; nine innings is a workout.