Papers for

machine learning platform operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Measuring debate shifts in AI model answers on multiple-choice tests

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Abstract: Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.

Mon 28 SeptComputation and Language
The gist
This paper looks at how AI models argue over multiple-choice questions and whether their discussions really help or hurt the final answer quality. The researchers find that just seeing the final result improves doesn't tell the full story, because discussions can either fix mistakes or cause new errors. They develop a detailed way to track these changes during debates, showing when answers collapse or get corrected. Their tools help figure out when debate helps and when it might backfire.
Open → 2609.35279v1

Neural network training verified cheaply without trusting trainers

Training Witnesses: Trusting the Training without Trusting the Trainer

Abstract: Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.

Sun 27 SeptMachine LearningCryptography and Security
The gist
Machine learning results depend on trusting the person who trained the model, which can lead to mistakes or unfair comparisons. The authors propose a way to check if training was done correctly without redoing the entire process. Their method uses quick behavioral checks and occasional tests to confirm the model's training and data use. This makes it easier and cheaper to verify models and helps create trusted benchmarks for progress.
Open → 2609.33915v1