Papers for

automated data labeling teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Smaller teachers improve knowledge distillation with less data

When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation

Abstract: Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.

Mon 28 SeptMachine Learning
The gist
Knowledge distillation is a way to train smaller, faster AI models by learning from bigger, more complex ones. This paper finds that when you have less training data, using a smaller teacher model can actually work better than a larger one. The authors explain this by looking at how teachers rank classes and score their predictions. They also show that choosing training examples that are just the right difficulty and offer varied information helps the smaller teachers do better. They propose a new method to pick these training samples without needing extra training details, which works well compared to other approaches.
Open → 2609.34489v1