Strategically sampled data improves training for complex AI tasks
Strategically Diverse Sampling for Self-Training
Computation and Language
Summary
Training AI models often involves using many examples that look very similar, which doesn’t always help them learn better. The authors studied how showing the AI examples that use very different problem-solving approaches, even if some are incorrect, can improve learning. They created new ways to pick these diverse examples and found that models trained this way perform better on hard problems. Surprisingly, this method can outperform training with much larger models when the data is chosen randomly.
What this means in practice
- •For machine learning engineers: Improve performance of AI models on difficult tasks by training with data sampled for diverse solving approaches rather than just correct answers.
- •For software developers: Enhance code generation AI tools by integrating training methods that include varied problem-solving paths rather than repetitive or similar code examples.
Authors
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Abstract
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.