Yesterday, my post about forcing ChatGPT, Claude, and Gemini into a roundtable discussion to fact-check eachother got way more traction than I expected.
The idea is simple: use the diversity of three AI models to catch hallucinations. If OpenAI misses a logical leap, Anthropic or Google catches it.
But some of the sharpest comments here pointed out the ultimate failure mode: What if all three models share the exact same training blind spot?
So instead of defending the setup, I want you to help me break it. Give me a question, problem or prompt that you think ChatGPT, Claude AND Gemini will all get wrong. It could be an obscure factual trap, a very convincing false premise, a common coding misconception, or a logic puzzle where the internet consensus is wrong.
The part I'm especially curious about is whether:
1. One model catches a mistake immediately
2. They fight and eventually figure it out
3. Or all three confidently agree on the same wrong answer
For context, this is the multi-model discussion setup I've been building into Rauno, but I'm mainly interested in finding its failure cases here.
Give me your best attempt on a question to break it and I'll reply if they actually caught each others hallucinations.
[link] [comments]