AI Operations · B10–B27wk 3 / 54
Writing

Longer pieces, when there's a real finding to write down

The working log captures what happened day to day. This is for the pieces that took longer to shape — usually because the result wasn't what I expected going in.

2026-09-10evaluationreliabilitypass@k

What pass@1 hides

A support-ticket classifier scored 66% on one run per case. Running each case 8 times and requiring all 8 to pass barely moved the number, which was the actual finding.