The first time I ran two LLMs against the same pull request, 89% of their "debate" was fake.
Not wrong. Not low-quality. Fake. The second model was replaying pre-generated text. Both models were exchanging messages, no positions were changing, no evidence was being cited — and the engine produced a confident verdict with a transcript full of sophisticated-sounding exchanges. If I hadn't dug into the raw logs I would have shipped it and called it a working system.
That's the problem I was actually building against — not "how do I get two models to review the same thing," but "how do I stop two models from performing agreement without ever genuinely challenging each other?" Because those are completely different problems. And almost every multi-model review pipeline I've seen solves the first while quietly ignoring the second.
The Structural Flaw Nobody Is Talking About
Here's what most "AI second opinion" workflows actually look like:






