Papers for

ai model training teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Representation alignment improves training speed of visual AI models

What Visual Generators Need from Teachers: Rethinking Representation Alignment

Abstract: Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet $256\times256$, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Training AI models that generate images is faster when parts of the model learn to copy specific layers of a well-trained reference model, called a teacher. The authors found that it only helps to copy teacher layers that the new model struggles to mimic on its own. They created a method called RARE that intelligently picks which teacher layer to copy and stops copying when improvements slow, saving time and producing better image quality. Their approach beats several older methods on standard benchmarks with less computing effort.
Open → 2609.34732v1

Self evolving language models learn faster without challenger training loop

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Abstract: Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.

Mon 28 SeptMachine LearningArtificial Intelligence
The gist
Language models can train themselves by creating tasks and learning from their own answers, but usually this needs a separate training step for the task generator. The authors propose a new method called Direct Self-Evolving Optimization (DEO) that removes this extra training step by letting the solver guide task selection directly. Their method trains only the solver model while using a fixed task generator and still improves reasoning skills efficiently. Experiments show DEO matches previous methods while using less training time and even benefits from using a frozen external language model to generate tasks.
Open → 2609.34279v1

Explicit preference modeling improves offline alignment of large language models

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

Abstract: Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.

Mon 7 SeptComputation and LanguageMachine Learning
The gist
Aligning large language models with human preferences is important but challenging. The authors study why methods that update models in steps usually work better than one-time offline methods. They find that having a clear model of preferences makes a big difference. Using this insight, they create a new method that learns this preference model first, then uses it to train the language model more efficiently. This new approach works as well as stepwise methods but takes less time to train.
Open → 2609.06893v1