Bayesian memory improves multi-agent advice reliability and task success
BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
Artificial Intelligence
Summary
Getting help from multiple agents can sometimes be tricky because not all helpers are good at every task, and some can even give misleading advice. The authors created BaRe-Mem, a memory system that learns how reliable each helper is by looking at past interactions and the main agent's own knowledge. This helps the main agent decide when to trust advice and when to act alone, leading to better results even when some advisors are wrong. They also showed BaRe-Mem can improve how teams assign tasks by picking workers who are more likely to succeed.
What this means in practice
- •For multi-agent system developers: Improve the reliability of agent consultations by estimating advisor trustworthiness from past interactions and internal beliefs.
- •For team task planners: Enhance worker assignment by dynamically identifying capable agents early based on historical success rather than simple success counts.
Authors
Peilin Feng, Zhengyang Huang, Soujanya Poria
Abstract
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.