Papers for

nlp engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Efficient attention method speeds up large context processing

SANTA++: Sampling Attention through Representative Keys

Abstract: Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a $1.69\times$ attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.

Mon 28 SeptMachine LearningComputation and Language
The gist
Processing long sequences of information in AI models can be slow because they have to look at everything at once. The authors introduce SANTA++, which cleverly picks smaller representative parts to pay attention to without having to check everything. This approach keeps results very close to full processing but reads much less data, making it faster and more efficient. Their tests show it works well on tasks requiring understanding of very long context.
Open → 2609.35629v1

Neural language models learn meanings from sentence structures not just words

Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality

Abstract: It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.

Mon 28 SeptComputation and Language
The gist
Language models usually learn by looking at word statistics, but it’s unclear how they understand sentence structures that carry meaning beyond words. The authors show that these models can learn patterns of how grammatical structures appear in different contexts, even when those patterns can't be guessed by looking at individual words. They trained models on a made-up language where each structure had unique context patterns and found that models first learn the grammar relationships and then the meanings from context. This suggests that models build understanding by treating grammatical structures as new units for learning meaning.
Open → 2609.34936v1

Large language models unevenly recognize names by race and gender

Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

Abstract: Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.

Mon 28 SeptComputation and LanguageArtificial IntelligenceComputers and Society
The gist
Different names aren't always treated equally by large language models because some names are recognized as single tokens, while others are broken into smaller parts. The authors show this unequal tokenization happens more often depending on the name's race or gender associations. They created a method called NameTrace to measure how this affects models’ understanding of concepts linked to those names. Results indicate that these differences matter for tasks like hiring or lending decisions made by the models.
Open → 2609.34065v1

Language models improve learning from context using perturbed documents

Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Abstract: Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.

Sun 27 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models often struggle to understand and reason with new information provided in a task's context. The authors demonstrate a method that alters public documents slightly to prevent models from just memorizing them and instead encourages true understanding. Using this approach, they generate synthetic training data without needing expensive human labeling. Training smaller models on this data helps them better use context for reasoning, improving their performance on certain benchmark tests.
Open → 2609.33642v1

Researchers map detailed internal features for time aware recall in language models

Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall

Abstract: Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing time-sensitive recall remain under-explored. Unlike prior studies that identify temporal components such as attention heads and MLP layers, we provide the first feature-level map of temporal recall by isolating individual MLP features via transcoder circuit tracing. We identify three node categories (common temporal, common to the year, and chrono-semantic) which interact to generate a temporal filter during factual recall. By analysing Gemma 2 2B, LLaMA 3.2 1B, and Qwen3-4B, we show that these features do not follow a simple linear pipeline but represent time through a parallel and mixed syntactic-semantic interplay across layers. We additionally discover a class of higher-layer temporal components invisible to existing EAP-IG methods, establishing transcoders as a more complete lens for temporal interpretability in time-sensitive factual recall. These findings present MLP components for potential targeted interventions in time-sensitive factual recall

Sun 27 SeptComputation and LanguageMachine Learning
The gist
Language models sometimes get facts wrong because they mix up info from different times. The authors found specific tiny parts inside these models that help them remember facts correctly based on time. They looked at different models and saw that remembering time happens in a complicated way, not just step by step. Their work can help improve how language models keep up with facts that change over time.
Open → 2609.33183v1

Large language models improve reasoning with semantic abstraction framework

Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models

Abstract: Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural language inference (NLI), which requires sophisticated linguistic capabilities to interpret implicit meanings, contextual conceptual relationships, and semantic connections between words and phrases. To this end, we propose a methodological framework for constructing new semantic knowledge at a higher level of abstraction, which we define under the notions of semantic compatibility and incompatibility for NLI. In this framework, the meaning of the lexical-semantic relations between the premise and the hypothesis is reconfigured to achieve a more flexible semantic network that induces different reasoning paths in LLMs. These new pathways show a consistent pattern of responses that allows agreement on a single response. The results demonstrate that our proposal allows to discover and compensate for LLMs' semantic knowledge gaps in NLI, achieving significant improvements in accuracy, exceeding 10% for some models, and in particular for the non-entailment class. It is essential to note that LLMs need structured knowledge and not just more data to bridge reasoning gaps. Our hybrid approach directs attention to overlooked word relationships, allowing models to synthesize missing information. We believe that the future lies not in increasing model size, but in creating a semantic scafolding that mimics the flexibility of human thinking. Hopefully, our proposal will enable the development of more robust agents and interpretable reasoning, guiding AI toward reliable language understanding.

Tue 22 SeptComputation and Language
The gist
Large language models (LLMs) sometimes struggle to understand deeper meanings between sentences during tasks like language inference. The authors propose a framework that reorganizes word relationships to help these models reason more like humans by discovering missing knowledge. Their approach boosts model accuracy by over 10%, especially in understanding when statements do not follow from each other. This method focuses on structured knowledge rather than just adding more data, helping LLMs fill in gaps in their reasoning.
Open → 2609.26610v1

Culture language and region annotations enhance web data benchmarking

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.

Mon 21 SeptComputation and Language
The gist
It’s hard to tell if language AI models understand different cultures because the training data and tests are labeled differently. The authors created FineWeb-CLaR, a big dataset that tags web documents with culture, language, and region information, making it easier to compare what’s in the training data with what tests measure. They also labeled many benchmarks to match this system, so it’s clearer if models truly cover cultural topics. This helps check if AI learns from or is tested on diverse cultural content.
Open → 2609.25298v1

Large language models show varied political leanings based on test setup

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

Abstract: Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.

Tue 8 SeptComputation and LanguageComputers and Society
The gist
Large language models are used to provide information, but it's tricky to know their political leanings because test results can be affected by how questions are asked and other factors. The authors created a detailed evaluation method that changes many test settings to see how these affect the models' political scores. They found that most models lean toward the Libertarian-Left, but changes in language, instructions, and answer formats can shift this. Also, smaller models can appear neutral simply because their answers are weak signals instead of truly centrist views.
Open → 2609.08637v1