For years, those of us who view testing as skilled knowledge work were told there’s no such thing as a “testing mindset.” We’ve also endured criticism that independent testing is just self-interested consultants preserving a business model.
Their argument was simple: developers can learn testing skills. Creation and critical evaluation don’t require fundamentally different mindsets. Anything other than that was a “myth” designed to create silos and stop knowledge from being distributed across teams and in fact, dedicated testing specialists might not be necessary at all.
Then AI became the developer and suddenly, everyone is worried about who evaluates the work.
And the research is explaining why LLM-as-a-Judge, generator/evaluator architectures, judge calibration and using multiple evaluators is the new “independent testing”.

This 2024 paper on LLM Evaluators found that LLM evaluators can “recognise and favour their own generations”. When the same LLM acts as both evaluator and evaluatee, self-preference can compromise what appears to be neutral evaluation.
And this research published in 2026 described “preference leakage” when generators and judges are the same model, inherit from one another, or belong to the same model family. They found judges systematically biased toward related models and the problem wasn’t bad prompting, but relatedness between creator and evaluator.
And to top it all off, this June 2026 survey of LLM-as-a-Judge identifies reliability and bias mitigation as fundamental problems in constructing trustworthy evaluation systems. Simply appointing an LLM as judge doesn’t make its judgment reliable.
I guess having the creator evaluate its own work might introduce blind spots after all.
It might sound petty (I don’t really care), but I can’t help remembering at the time this mythology nonsense was being pushed that James Bach disputed the straw man directly when it was published. Years later, AI researchers may not be using his terminology, but they are describing the same problem.
And credit where credit is due, my friend Michael Bolton has been making a similar argument for nearly twenty years: the problem isn’t whether a developer can test, it’s whether a creator can achieve sufficient critical distance from their own work.
But as with all things in the era of Testing with AI, everything old has become new again as we watch the AI Testing Bros discover skilled testing. “Critical distance” has become “evaluator independence”. “Builders perspective” has become “generator”. “Tester perspective” has become “judge/evaluator”. “Mindset” has become “self-evaluation bias”, and on and on it goes.
But until they find a Testing Mindset neuron inside Claude, the value of independent testing has always been the same: someone deliberately occupying a different epistemic position, making different assumptions, and using different models of risk and purpose.
And that’s not a myth.
Discover more from Quality Remarks
Subscribe to get the latest posts sent to your email.