Agent Evals · live quality record

Three shipped agent builds.
Here is what they actually score.

Most agent portfolios show you the happy path. This one runs a frozen golden set against three live systems, scores every output against anchored rubrics, and keeps the history — including the failures. Nothing here is mocked or replayed. Every number is whatever the real endpoints returned on the date shown.

Latest overall
Cases in set
Checks failing
Weakest dimension
Recorded runs

Loading stored history…

Latest run

Live runs spend real Workers AI neurons, so public runs are limited to one per 10 minutes.

Score drift over time

One score is a screenshot. A series is evidence. Same inputs every run, so a drop here means the system changed — not the test.

How it scores

Frozen golden set

Inputs never change (). Two cases are adversarial: one plants a fabricated claim the verifier must reject, one starves the planner of context to see whether it invents facts.

Deterministic checks carry 60%

JSON contract, array bounds, near-duplicate detection, latency budget, refusal discipline. No model involved, so these cannot drift.

Anchored judging carries 40%

Each dimension ships explicit 1/3/5 descriptors. Without anchors an LLM judge returns 4 for everything. Judge model is pinned: .

Evidence or the score is void

Every judged score must quote the output verbatim. The quote is checked against the real output in code. A quote that isn't there voids the score.