Networks with two parents per agent achieve exact information aggregation
Optimal Networks for Agentic Information Aggregation
Machine LearningComputer Science and Game Theory
Summary
This paper looks at how groups of agents (or computers) can work together to learn from pieces of information they each see. The authors study how the way agents connect and share predictions affects how well they can match the best overall prediction using all information. They find that if each agent only gets one parent's input, it’s impossible to perfectly match the best prediction. But if each agent gets two parents’ inputs, then it is possible, even if the network design doesn’t know the information ahead of time. They also find efficient ways to build such networks with reasonable size and depth.
What this means in practice
- •For distributed system builders: Design communication networks that exactly combine distributed data sources into a best linear prediction using minimal connections per node.
- •For machine learning engineers: Construct efficient networked models with limited node inputs that still achieve optimal prediction accuracy across features without tuning to specific data distributions.
A theory result. No direct application yet.
Authors
MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Shayan Taherijam
Abstract
We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents' predictions, fits a linear predictor to minimize mean squared error, and passes only its prediction forward. The global predictor is the best linear predictor using all features. Kearns, Roth, and Ryu show that the output agent's error approaches the global predictor's error along sufficiently deep paths with suitable feature coverage, while insufficient depth can prevent aggregation even in large networks. In contrast to their main focus on a given graph and feature allocation, we consider the limits of the model under two settings. In the adaptive designer setting, a designer chooses the graph, feature allocation, and output agent knowing the distribution. In the oblivious designer setting, the designer fixes all three before an adversary chooses the distribution. Each agent observes one feature and receives predictions from a limited number of parents. We call the aggregation exact when the output agent matches the global predictor exactly. For $d\ge3$, we show that no finite depth guarantees exact aggregation for every distribution with one parent per agent, even when the designer knows the distribution. In contrast, two parents per agent suffice for exact aggregation even in the oblivious designer setting. A fixed graph, feature allocation, and output agent achieve this for every distribution at depth $O(d\log d)$. Knowing the distribution reduces the depth to $O(d)$. Both constructions use $O(d^2)$ agents, with a very large constant for two parents. We show the bounds on the depth and number of agents are all optimal up to constant factors.