B10 · measurement · Live operation
Evaluation Harness. The Foundation
Why this block exists
Scheduled here because it comes it has zero upstream dependencies and everything measured before it exists is retrofitted and unverifiable. This was the single biggest correction in the plan's design.
What it has to produce
Build the measurement layer that decides whether any agent you ever ship actually works. Ship it as an installable package. Establish the golden-set discipline every later block depends on.
01What it has to teach me
topicpass^k reliability methodology
topicLLM-as-judge validation
topicgolden set design
topicconfidence intervals
topicagent interface contracts
02Evidence this page will carry
emptyPrimary artifact
emptyMeasured result
emptyRejection record
2026-09-01Outline created from the AI Operations Roadmap (B10-B27).