Deliberation among diverse ai models improves collective accuracy

The Wisdom of Artificial Deliberative Crowds

Artificial Intelligence

Summary

Sometimes, a group of people working together can make better guesses than experts alone by talking things through. This study tested if the same idea works with different AI models talking to each other. They found that when diverse AI models discuss and agree, their group answers are better than just averaging individual predictions. Also, each AI model becomes better on its own after this group talk. But if the group is made up of copies of the same AI, the benefit disappears.

What this means in practice

  • For machine learning teams: Improve group decisions by combining diverse AI models that deliberate to enhance the accuracy of peer reviews and AI behavior detection.
  • For sports analysts: Use ensemble AI deliberation approaches to boost the accuracy of sports game outcome forecasts beyond traditional prediction markets.

Authors

Federico Barrera-Lemarchand, Mariano Sigman, Joaquin Navajas

Abstract

The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation. Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4). Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.