Most agent portfolios show you the happy path. This one runs a frozen golden set against three live systems, scores every output against anchored rubrics, and keeps the history — including the failures. Nothing here is mocked or replayed. Every number is whatever the real endpoints returned on the date shown.
Loading stored history…
One score is a screenshot. A series is evidence. Same inputs every run, so a drop here means the system changed — not the test.
Inputs never change (—). Two cases are adversarial: one plants a fabricated
claim the verifier must reject, one starves the planner of context to see whether it invents facts.
JSON contract, array bounds, near-duplicate detection, latency budget, refusal discipline. No model involved, so these cannot drift.
Each dimension ships explicit 1/3/5 descriptors. Without anchors an LLM judge returns 4 for
everything. Judge model is pinned: —.
Every judged score must quote the output verbatim. The quote is checked against the real output in code. A quote that isn't there voids the score.