Airport scanners show screeners fake guns on purpose. It's the only way to know they still look. And the better your AI agent gets, the worse your reviewer gets at catching it.
Last week I asked one question about any process an agent already runs for you: when did the reviewer last say no? (https://lnkd.in/p/drJzefVy) The honest answer for most teams is "I don't know," and the hopeful one is "never, which is good, right?" This piece is the follow-up. It's longer than a post because the answer is arithmetic, and I'd rather show the arithmetic than ask you to trust me. Every formula fits on one line, and every one comes with an example.
One number, two unknowns
A rejection happens when two things happen together. The agent made a mistake, and the reviewer noticed it. Call the first the error rate, p. Call the second the catch rate, c. The only thing you can see from the outside is the rejection rate, and it's roughly their product:
rejection rate ≈ p × c — how often the agent is wrong, times how often the reviewer notices.
A perfect agent with a sharp reviewer gives you zero rejections. A sloppy agent with a reviewer who signs everything also gives you zero rejections. The dashboards look the same. You can collect data for ten years and they will keep looking the same, because more data only pins down the product more precisely. It never splits it.
That's the whole problem. Everything below is a way to break the product apart.
If nothing goes wrong, is everything all right?
That's the title of a 1983 paper in JAMA by James Hanley and Abby Lippman-Hand, and I can't improve on it as a heading. The paper gives the rule of three: if something hasn't happened once in n tries, you can say with 95% confidence that its rate is below 3/n.
zero events in n tries → the rate is probably below 3 ÷ n.
Say a reviewer has signed 500 agent outputs with zero rejections. The rule of three says p × c is below 0.6%. Now bring in what you think you know about the agent. If its error rate is 2%, the reviewer catches less than a third of its mistakes. If the reviewer really catches everything, the agent is wrong less than 0.6% of the time, and you should ask the vendor why their own number is three times worse.
Turn it around. Suppose the agent really is wrong 2% of the time and the reviewer really catches every mistake. The chance of seeing 500 clean signatures in a row is 0.98 to the power of 500, about one in 24,000. A long quiet streak isn't proof of anything bad. It just isn't the good news it looks like.
Plant the mistakes
The way to split p from c is to stop waiting for mistakes and make some.
Airport security has done this for a long time. The system is called Threat Image Projection. The X-ray machine inserts a picture of a gun or a knife into a real bag on the screen. If the screener flags it, the machine says "that one was fake, carry on." If they don't, it goes in the record. In the US the FAA started rolling it out in 2000, with a library of about 2,400 threat images. Software people had the same idea earlier: Gerald Weinberg suggested planting bugs in 1970, and Harlan Mills at IBM turned it into an estimate in 1972.
The arithmetic is short. You plant m known mistakes and the reviewer catches k of them. Your estimate of the catch rate is k ÷ m. Once you have c, the real rejections tell you how many real mistakes there were.
catch rate ≈ caught ÷ planted; real mistakes ≈ real rejections ÷ catch rate.
Example. You plant 20 mistakes over a month and the reviewer catches 12, so c is about 60%. The same month there were 30 real rejections. That means about 50 real mistakes, and about 20 of them went out the door with a signature on them.
Nobody should get fired over this, by the way. You're measuring what the signature is worth, and the reviewer is the instrument, not the suspect.
How many to plant
Twelve out of 20 sounds like a measurement. It mostly isn't. With 20 seeds the 95% range for the catch rate runs from about 40% to about 80%. That's the difference between "the reviewer is fine" and "the reviewer misses most of it," in one interval.
To know the catch rate within ten points either way, you need about 100 planted mistakes in the period you're measuring. If the team signs 500 outputs a month, that's 20% extra work every month, which nobody will approve. So measure per quarter, or across the whole team rather than per person: 100 seeds spread over 1,500 signatures is under 7%.
The conclusion I didn't expect: planting has to keep running in the background, like fire drills in an office building. Run it once, get a number, stop, and in six months you're back to staring at a product of two unknowns.
The seeds have to look real
All of the above rests on one assumption: the planted mistakes are as hard to spot as the real ones. If they're easier, the catch rate comes out too high and everything built on it is too optimistic.
Ecologists have a word for the version of this that bites them: trap-happy. Some animals learn that the trap has food in it and walk into it again and again. They get recounted, and the population looks smaller than it is. A planted mistake that's obviously planted (a total of $1,000,000 on a $40 invoice, a candidate named Test Testerson) is the trap-happy animal. The reviewer catches it every time and learns nothing about the reviewer.
The best seed is a real mistake from the past. If you keep a log of what reviewers corrected in agent drafts, you already have a warehouse of them: realistic, from your own domain, in the agent's own style. Take one from last spring, put it into this week's queue, see if it gets caught again.
Two reviewers and no planting
If planting feels too artificial, there's an older trick that needs no fake mistakes at all, only a second pair of eyes. It comes from counting fish. Carl Petersen used it on young plaice in Denmark in 1896, and Frederick Lincoln on ducks in 1930. Catch some animals, tag them, let them go. Catch again later. The share of tagged animals in the second catch tells you how big the whole population is. In 1992 Stephen Eick and colleagues at Bell Labs applied it to software design reviews.
Swap animals for mistakes and catches for reviewers. Two reviewers check the same batch independently.
total mistakes ≈ (found by A × found by B) ÷ found by both.
Example. Reviewer A finds 40 problems, reviewer B finds 30, and 20 of them are the same. The estimate is 40 × 30 ÷ 20 = 60 mistakes in the batch. Together they found 50, so about 10 nobody saw.
The catch is in the word independently. Two reviewers trained by the same person, using the same checklist, miss the same things. So do two models from the same family. When their misses overlap, the "found by both" number goes up and the estimate of the total goes down. Two reviewers who each catch 90% miss 1% together if they're independent, and 10% if they're copies of each other. The formula can't tell which case you're in. That deserves its own piece, and I'll come back to it.
The better the agent, the worse the check
This is the part that changed how I think about the whole thing.
In 2005 Jeremy Wolfe, Todd Horowitz and Naomi Kenner published a short paper in Nature called "Rare items often missed in visual searches." People looked at baggage-like images for a target. When the target was in half the images, they missed 7% of them. When it was in 1%, they missed 30%. Same people, same task, same targets. The only thing that changed was how often there was anything to find.
Now apply that to agents. Every time your agent gets better, its mistakes get rarer, and the reviewer's eye gets less practice at finding them. Improving the agent makes the check worse on its own, without anyone getting lazy. A reviewer who used to see a bad draft every day now sees one a month, and the one a month gets through.
Planting helps here twice. It measures the catch rate, and it raises how often there's something to find. Plant 100 seeds into 500 outputs where the agent makes 10 real mistakes, and the reviewer sees a problem in about one output in six instead of one in fifty. That won't cure it. It does push the job back toward the conditions where people are good at it.
Where you don't need any of this
In the oversight post I argued that the useful question about a step isn't whether you can undo it but how soon you find out it was wrong. The same line decides where planting is worth it.
Where you find out right away, reality plants the mistakes for you. The payment bounces, the build goes red, the customer calls the next morning. You see the agent's errors whether the reviewer caught them or not, so you can compute the catch rate from what slipped through.
Where the answer is "never" or "a year later, in a complaint to the regulator," nothing comes back on its own. A refused claim, a rejected candidate, a contract clause nobody reads until the dispute. That's where the reviewer is the only line, and that's where you have to measure the line.
What the vendor's number is worth after a quiet month
The last tool is the one I'd actually put on a dashboard.
Start with what you believe. Say the vendor tells you the agent is wrong 2% of the time, and you think the reviewer is fine: catches about 80%, and you're 90% sure of that. The other 10% is the worry that the reviewer has quietly become a rubber stamp who catches 10%.
Then 200 outputs go by without a single rejection. If the reviewer is fine, the chance of that is about 4%. If the reviewer is a rubber stamp, it's about 67%. Put those against your starting 90/10 and your confidence in the reviewer drops from 90% to about 35%. Nothing visible happened. Two hundred signatures, no drama, and the sensible belief flipped.
You can turn that into a rule anyone on the team can apply without a calculator:
a run of zero rejections longer than 3 ÷ (p × c) is a reason to check the reviewer.
With a 2% error rate and an 80% catch rate, that's about 190 outputs. A streak that long happens by chance roughly one time in twenty. It proves nothing. It tells you to go and look.
What I'd do on Monday
Pick the process where an agent's mistake would surface latest. Count the signatures since the last rejection and compare that with 3 ÷ (p × c). Pull a dozen real mistakes from the correction history and slip them back into the queue over the next few weeks. Where you have two reviewers who don't share a checklist, have them check the same batch once and compare. And whatever number you get, plan to get it again next quarter, because the better the agent gets, the more that number drifts.
None of this needs new software. It needs someone to decide that the reviewer is part of the system, so it gets measured like the rest of it.
For the process you trust the agent with most: how many of its mistakes has your reviewer caught that you put there on purpose?
— Maxim Berg, Founder, HRPulsar