AI Operations · B10–B27wk 2 / 54
B10 · measurement · Live operation

Evaluation Harness. The Foundation

Status
planned
Weeks
13
Permission gate
none
Depends on
nothing
Page depth
outline
This is an outline, not a result.
Nothing here has been built. It states what the block is for, what it has to produce and what would make it a failure, so that the claim can be checked against the result later. Written 2026-09-01, and it will change.

Why this block exists

Scheduled here because it comes it has zero upstream dependencies and everything measured before it exists is retrofitted and unverifiable. This was the single biggest correction in the plan's design.

What it has to produce

Build the measurement layer that decides whether any agent you ever ship actually works. Ship it as an installable package. Establish the golden-set discipline every later block depends on.

01What it has to teach me

topicpass^k reliability methodology
topicLLM-as-judge validation
topicgolden set design
topicconfidence intervals
topicagent interface contracts

02Evidence this page will carry

emptyPrimary artifact
emptyMeasured result
emptyRejection record
2026-09-01Outline created from the AI Operations Roadmap (B10-B27).
B8 Smart Reply Copilot: Multi-Step Prompt ChainingB11 Surface Zero Goes Live