Multi-agent debate (MAD) has long been reported to improve reasoning and factuality over single-model inference. However, a new research paper published on arXiv challenges the status quo, investigating whether "cognitive diversity" among agents is truly the driver behind these perceived gains.

Rejecting the Diversity Hypothesis

Researchers tested 23 open-weight models from eleven vendor families across five tasks and over 5,500 runs. By varying diversity along three axes—personas, sampling temperature, and model identity—they compared every debate configuration against a generation-budget-matched majority-vote control. The hypothesis that diversity drives performance was rejected on every axis.

While debate does outperform single-agent inference by 3 to 7 points in certain tasks, it fails to justify its cost. Under budget-matched conditions, debate ties or even loses to self-consistency sampling, despite being 1.6$ imes$ slower in wall-clock time and 3.4$ imes$ more expensive in terms of token consumption.

The 'Persona Tax' and Context Overflows

The study identifies a "persona tax," where persona prompting actually reduces accuracy. Redundant personas were found to hurt performance the most, while maximally diverse teams only managed to recover a portion of that loss. Furthermore, the accuracy of mixed-model teams tracked the capability of individual members rather than the heterogeneity of the group.

A significant technical discovery involved a "pervasive measurement hazard." Debate transcripts frequently overflow serving context windows without warning. Correcting for this overflow alone shifted the comparison between debate and sampling from a 1.8-point deficit to parity.

Redefining Multi-Agent Gains

The researchers found that nearly all the benefits of debate occur during the very first exchange of answers. Ultimately, the results recast reported MAD gains as an ensemble-sampling effect rather than a product of meaningful cognitive interaction. The study establishes a new baseline, suggesting that future debate mechanisms must clear the bar of budget-matched, contamination-checked majority voting to be considered effective.