- date
- words
- 4 749
- read
- 24 min
- sources
- 7 sources
- scenarios
- 41 scenarios
- runs
- 123 runs
- suite cost
- $3.06 suite cost
- undercount
- 4.5x undercount
- untested tools
- 6/17 untested tools
- views
- 16 views
Six of our agent's seventeen tools had never run.
Six of seventeen agent tools had never once run in production, including both of the ones that unlock a contact and charge for it. This is the harness that finally tested them — a real model in a completely faked world, 41 scenarios, 123 runs, $3.06 — and the cost blind spot it uncovered on the way.
AI Agents · Evaluation · LLM