Confidently Incorrect

You want to know why you shouldn’t try to “engineer” confidence? Check out the latest incident report from Anthropic where “given the volume of transcripts and our desire to disclose incidents quickly”, they used an AI-powered inspect-o-nator to review over 140k transcripts and STILL missed the exact behavior they were trying to detect. 

Claude’s own reasoning convinced the evaluator that the environment was simulated. With that reasoning included, the monitor flagged about 1% of its actions but if you take Claude’s explanation away and show the evaluator the actual tool calls and results? About 50%. The checker got worse because it trusted the thing it was checking!

For those of you in the back and all the AI Test Bros discovering testing: checks are not testing!

Michael Bolton and James Bach have been talking about this for years. A check can tell you whether an observation matches a predefined rule. Testing asks if you looked in the right place, can you even observe the thing your looking for, can you trust your oracles , what assumptions are you making and what’s the risk to our business if we get it wrong.

And if you’re still thinking that the systems being built/enhanced with AI coding agents are somehow being governance by human agency, you’re not paying attention. Verification Asymmetry is an intractable problem with GenAI extruded code and the testing industry response (and pretty much the agentic world as well) is to just throw more AI generated stuff at it. Now we get to see the consequences of that in real time. 

All while the algorithm spits out some number telling us how confident we should feel…


Discover more from Quality Remarks

Subscribe to get the latest posts sent to your email.

Leave a Reply