Papers for

machine learning teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Distributed teams optimize complex tasks by sharing compact summaries

GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization

Abstract: Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO, in which agents exchange compact distributions over the locations of their respective optima inferred from local Gaussian process (GP) posteriors, rather than raw observations, query points, or surrogate parameters. The server merges and reweights these distributional components before returning a subset to each agent. Each agent then constructs a Federated Interventional GP (FI-GP), which preserves the local posterior mean and spatially rescales its covariance for local decision making. For the upper confidence bound (UCB) instantiation, GUIDE-UCB, we prove that any bounded FI-GP uncertainty intervention preserves the leading-order cumulative regret rate of standard GP-UCB. When the transferred distributions place greater support near an optimum than in a suboptimal region, selecting the latter requires greater local posterior uncertainty. Experiments on 12 synthetic benchmarks and three real-world optimization tasks show that GUIDE-FBO remains effective across settings ranging from homogeneous to severely heterogeneous. Ablation results highlight the importance of spatially localized uncertainty intervention, while the communication analysis shows that GUIDE-FBO exchanges only compact distributional messages.

Mon 28 SeptMachine Learning
The gist
Optimizing costly and unknown processes often requires multiple teams working together without sharing sensitive data. The authors present a method where teams exchange small summaries describing their best guesses about the solution instead of raw data. These summaries help each team improve its own decisions while respecting communication limits and differences in tasks. Their approach performs well on both simulated and real-world problems, maintaining efficiency despite varying team needs.
Open → 2609.35038v1

Machine unlearning aims to remove specific influences in large language models

Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions

Abstract: Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.

Fri 25 SeptCryptography and Security
The gist
Sometimes developers want to make large language models forget specific information without hurting the model's other skills. This paper looks at different ways to do this, compares tools and tests, and suggests a framework to analyze how well unlearning works. The authors find current tests are not enough to prove that the unwanted data is truly removed, especially when models update or work with extra tools. They suggest better evaluation methods to track if the forgotten information can come back.
Open → 2609.30909v1

Deliberation among diverse ai models improves collective accuracy

The Wisdom of Artificial Deliberative Crowds

Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation. Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4). Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.

Fri 18 SeptArtificial Intelligence
The gist
Sometimes, a group of people working together can make better guesses than experts alone by talking things through. This study tested if the same idea works with different AI models talking to each other. They found that when diverse AI models discuss and agree, their group answers are better than just averaging individual predictions. Also, each AI model becomes better on its own after this group talk. But if the group is made up of copies of the same AI, the benefit disappears.
Open → 2609.22497v1

Consensus federated learning matches centralized training for robot vision models

Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models

Abstract: Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings.

Thu 17 SeptRobotics
The gist
Training robots to understand images, language, and actions often needs lots of data collected in many places, which is hard to gather in one spot. The authors developed Co-VLA, a method that lets many robots train a shared brain without sharing their private data by reaching agreement through a consensus process. This approach works well whether training the entire model or fine-tuning parts of it, performing as good as if all data were gathered centrally. Their method helps overcome challenges of different robots having different types of data.
Open → 2609.19923v1

ScientistTwo AI autonomously solves complex science problems

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Abstract: Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/

Thu 17 SeptArtificial Intelligence
The gist
Solving new science problems means exploring what we don't know yet. The researchers introduce ScientistTwo, a smart AI that tackles scientific questions all by itself. It tests ideas, runs experiments, and even reviews its own work without people helping. Their tests show ScientistTwo writes expert-quality papers and produces working code that often beats human results. This AI acts like a scientific explorer, pushing knowledge forward without needing human guidance.
Open → 2609.19644v1

Comparing retraining methods reduces error gaps over deployment time

Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

Abstract: Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affect these comparisons. In a follow-up sample of 400 new trajectories per condition across two simulated drift regimes, all three policies had lower mean cumulative disparity, equivalent to reductions of 0.04 to 0.88 percentage points in the average gap per window. Evaluating the unchanged models against the known generating distributions preserved all mean directions, but finite-window and population comparisons agreed on whether updating increased, reduced or left cumulative disparity unchanged in 69 to 92 percent of trajectories. Under subgroup-specific drift, smaller true-positive-rate gaps accompanied lower sensitivity in both groups. In an exploratory American Community Survey replay, person weighting reversed all three race false-positive-rate mean comparisons without changing predictions or actions; all three weighted intervals included zero. Policy comparisons require group-specific rates, action distributions, and an explicit evaluation population alongside mean disparity. These analyses are non-confirmatory. Shared replay requires policy-independent observations and complete labels after the specified delay.

Wed 9 SeptMachine Learning
The gist
Deciding when to update a computer model that makes decisions can affect how fairly it treats different groups. The authors compare several ways to decide when to retrain models based on performance changes or fixed schedules. They find that all updating methods generally reduce differences in error rates between groups compared to keeping the original model. However, the best method depends on factors like how errors appear over time and how groups change. Their work helps better evaluate fairness over a model’s lifetime.
Open → 2609.09788v1