When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn’t the only failure mode.
Some context. I've been running setups where a few LLM personas debate a question, then a separate neutral pass pulls out where they actually disagree. The whole reason I started was sycophancy. One model on its own just agrees with whatever you sa…