Papers for

data pipeline developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ReScraper improves web data cleaning for better large language models

ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

Abstract: LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language model of only 0.6B parameters that replaces this entire stack. To train ReScraper, we carefully curate supervised data from the outputs of three teacher models, so it learns to first extract the main content from raw data and then choose among four operations: keeping the page as extracted, editing out noisy lines and spans, deleting it entirely, or rewriting it when it is poorly written but informative. Based on the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on our curated data improves the DCLM Core score by a relative 3.8--4.7% over the strongest baseline at each scale, including the costly multi-agent curation. Our analyses show that each operation plays a distinct and complementary role, and that extracting and cleaning in one model outperforms a cascade of separate models. ReScraper also concentrates its operations on the pages that need them, raising the quality of poor pages the most while keeping the corpus diverse. These results demonstrate the feasibility and effectiveness of AI4AI for pretraining data curation, where a small learned model takes over an entire stage of the pipeline from hand-written heuristics. We open-source our code at https://github.com/cxcscmu/ReScraper

Mon 28 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Training large language models requires lots of clean text from the internet, but current cleaning uses many hand-coded rules that can miss problems. The authors created ReScraper, a small AI model that learns to find the main content on web pages and fix or remove messy parts automatically. This unified approach improves the quality of the training data more effectively than previous multi-step methods. Their tests show smarter cleaning helps models learn better language skills without losing variety in the data.
Open → 2609.34287v1

Coreset selection methods often cost more than they save in training

Are Coreset Selection Methods Worth Their Cost?

Abstract: Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.

Sat 19 SeptMachine LearningArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Coreset selection tries to speed up training by picking a small, representative part of the data. The authors tested many selection methods fairly by counting both the time to choose data and to train models. They found that simple random selection or using all data for fewer training steps often works better than complex selection methods. The cost to pick data is not small and often outweighs the benefit of training on fewer points. Ignoring the selection cost gives a misleading idea of how efficient these methods really are.
Open → 2609.22894v1

How graph design controls gradient flow in sparse sinkhorn layers

Support Topology and Gradient Mixing in Sinkhorn Layers

Abstract: Sparse Sinkhorn layers use a fixed support graph to restrict transport between tokens. How does this graph control gradient propagation through the scaling iterations. We develop a fixed-support calculus showing that each row-column cycle induces a row-stochastic operator on column-potential perturbations modulo constants. Its transpose propagates zero-mass reverse-mode cotangents. The finite-cycle operator uses two distinct half-step transport plans; at a balanced fixed point it reduces to a two-step walk determined by a single plan. We derive the accompanying score and marginal source terms and use Dobrushin contraction and minorization to bound homogeneous and source-driven tail cotangents. Our main result characterizes when support and marginals guarantee one-step contraction uniformly over finite scores: every feasible face of the transportation polytope must have pairwise two-hop column overlap. Otherwise, suitable score directions make the contraction coefficient arbitrarily close to one. We extend this analysis to ordered support schedules and derive certificates for partition heat-bath layers, coordinate sweeps, forced shared mass, and register-augmented supports. These results provide mathematical criteria for support design in differentiable transport layers, with guarantees restricted to the fixed-support quotient-gradient component.

Mon 7 SeptArtificial IntelligenceMachine Learning
The gist
This paper studies how the structure of a fixed support graph affects the way gradients pass through a special mathematical layer called a sparse Sinkhorn layer, used in optimal transport problems. The authors develop a mathematical approach to understand how changes in one part of the graph influence others during training. They identify precise conditions on the graph that ensure stable and consistent gradient flow, which is important for reliable model training. Their results provide guidelines to design these graphs so that training behaves predictably.
Open → 2609.07954v1