I’ve written in the past about the path we’re headed down in the software testing and quality engineering business. Expectations are being lowered to the point where AI generated and executed checks will be the primary method for testing systems. This will be paired with the “hands thrown up in the air” approach of the agentic coding community as they concede to Verification Asymmetry and just give up on reviewing AI generated code.
This phenomenon I refer to as “The Great Capitulation”, is no more apparent than in a recent article published by Fast Company called “Software That Can Check Its Own Quality”.
Their point seems to be that AI is generating software faster than traditional testing can keep up, so the only answer is autonomous AI testing. AI generates the code, AI generates the tests, AI maintains those tests, AI analyses the results, and humans move into the higher-order role of overseeing the STLC. Eventually, according to the author, “AI writes code, AI tests code, and humans are in the loop.”

The problem is that when their “quality loop closes”, the human is in the loop like a sock in a dryer.
A few days after the Fast Company piece appeared, a former OpenAI safety researcher wrote in The New York Times about AI systems whose developers cannot predict what boundaries they’ll obey. Anthropic then published its own analysis of cybersecurity incidents involving its models and acknowledged that creating evaluations which genuinely represent what those systems will do in real-world deployment remains an open research problem.
Both should make anyone advocating for autonomous agentic testing rethink that position.
Anthropic is discovering that AI evaluators can be influenced by the reasoning of the systems they evaluate, and European regulators are explicitly warning about automation bias and requiring competent human oversight capable of challenging the output of these systems.
The EU AI Act‘s approach to human oversight is remarkably different from our industry’s casual references to “human in the loop” (HITL). Article 14 requires high-risk AI systems to be designed and developed so that they can be effectively overseen by natural persons during their use and Article 26 places obligations on deployers to assign oversight to people with the necessary competence, training and authority to do just that.
The Act says those overseeing high-risk systems must be able to understand their capabilities and limitations, monitor for unexpected behavior, remain conscious of the risk of automation bias, and be capable of disregarding, overriding or reversing the system’s output and, where appropriate, stopping it altogether.
And that’s where all this agentic testing and evaluation create yet another closed loop. The more humans depend on AI to decide which evidence matters, the less independently capable they are of evaluating the system they are charged with verifying and HITL just becomes acceptance.
The regulatory model of assurance seems very different than an enormous pile of passing checks furiously clicked “approved” by middle management to build a workflow where “coverage scales automatically”.
Amazingly, the same vendor claiming software can now “check its own quality” appeared to acknowledge how problematic that statement is in another article published just six days later, stating that “a system cannot independently confirm its own reading of a requirement.” No kidding. Their answer is – you guessed it – ANOTHER AI loop acting as the independent verifier.
The paper the second article cites doesn’t even demonstrate that autonomous agents can independently establish software quality. Its focus is much narrower: agents work best when we constrain them to problems where someone has already made the answer objectively checkable. That’s not evidence that AI has solved testing or can check its own quality. It’s evidence that AI is getting very good at checking.
And in a spectacular display of irony, when the paper’s own authors needed to establish whether their multi-agent system had produced reliable results, they didn’t ask another agent and call the loop closed. They manually verified the output. Go figure.
Vendor claims in the software testing world have always had a hard time standing up to scrutiny and the word “coverage” has been doing Trojan amounts of work for decades. So statements in the article like “the productivity gains of autonomous testing are self-evident” should be met with a great deal of skepticism in a business that is ironically founded on the idea that NOTHING is self-evident.
Otherwise, we find ourselves inside another closed loop where a vendor defines the problem, provides the solution, provides the evidence, evaluates the evidence, claims the benefits are self-evident and we can all point to the article as independent validation of the trend.
Apparently, AI isn’t the only thing capable of hallucinating.
It can happen between vendors, analysts, conferences, executive networks and customers until repetition begins to look like evidence. At that point skepticism gets rejected rather than the basic intellectual rigor we should demand from anyone making claims about quality and testing.
No. Software cannot “check its own quality”.
So maybe the problem isn’t confined to using one AI loop after another to verify its own output, but another loop detached from reality operating alongside it: the software testing echo chamber surrounding autonomous testing.
Discover more from Quality Remarks
Subscribe to get the latest posts sent to your email.