Notes from the workshop

What we are building, what broke, and what the numbers said. Written by hand, published when there is something worth saying — not on a schedule.

tagged AI Agentsshow all

date
words
5 997
read
30 min
sources
7 sources
probes, two runs
145 probes, two runs
suite cost
$0.105 suite cost
identical replies
99/114 identical replies
distinct replies
10 distinct replies
indirect leaks
0/28 indirect leaks
cost doc drift
33x cost doc drift
views
8 views

Our agent passed every red team probe. That was the problem.

We pointed a generated red team at our agent and it passed everything. Then we counted the replies: 99 of 114 were byte-identical. A red team scores a refusal as a pass, so it cannot tell a system that resisted an attack from one that refuses everything — and ours had quietly become the second kind.

AI Agents · Evaluation · Security

date
words
4 749
read
24 min
sources
7 sources
scenarios
41 scenarios
runs
123 runs
suite cost
$3.06 suite cost
undercount
4.5x undercount
untested tools
6/17 untested tools
views
16 views

Six of our agent's seventeen tools had never run.

Six of seventeen agent tools had never once run in production, including both of the ones that unlock a contact and charge for it. This is the harness that finally tested them — a real model in a completely faked world, 41 scenarios, 123 runs, $3.06 — and the cost blind spot it uncovered on the way.

AI Agents · Evaluation · LLM