Papers for

computational linguists

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multi agent workflows learn when to skip steps for efficiency

Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

Abstract: Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.

Fri 25 SeptArtificial Intelligence
The gist
Large language model workflows often use several steps like planning, checking, and summarizing to get better answers. But running all these steps every time can waste work or even mess up good results. The authors created a method called Learning What to Skip (LW2S) that figures out when some steps can be safely left out by learning from past runs. This helps save computing resources while keeping or improving accuracy across tasks like math problems and coding.
Open → 2609.30734v1

Large language models excel generating sentences but struggle parsing semantics

Scoring Both Directions: LLMs realize the MRS they cannot reliably parse

Abstract: The English Resource Grammar (ERG) is a hand-written computational grammar of English. Given a sentence, its processor, ACE, produces a formal meaning representation called Minimal Recursion Semantics (MRS): a graph of the sentence's predicates and their arguments. The grammar is bidirectional and can also turn an MRS back into an English sentence. \citet{hajdik2019} used the ERG's treebank to build a benchmark for that generation task, MRS to text, and trained sequence-to-sequence models to solve it. The parsing task, text to MRS, can be tested on the same sentences. We reconstruct their 10K-sentence test split, and score two large language models, Claude Sonnet~4.5 and Claude Opus~5, in both directions against their trained systems and against ACE, with no task-specific training. Given an MRS and three examples, Opus writes the sentence at 76.3 BLEU, ten points above their system trained on 72k pairs (66.1 BLEU), and comparable to their system trained on a million extra pairs (77.2 BLEU). Sonnet scores 65.7 BLEU, and letting it choose among ACE's own candidate sentences lifts it to 69.6, while a pooled judge that keeps Opus's own sentence among the candidates adds 0.6 points (77.0 BLEU). In the parsing direction, however, the models fall far behind ACE: asked for the MRS of the same sentences, they reach 57.2 (Sonnet) and 65.5 (Opus) F$_1$ on the graph's predicates and arguments against 91.0 for ACE, and exact-match the gold on about 1\% of sentences. We characterize the failure modes for the parsing tasks, and conclude that a generation score alone does not show that models understand formal semantic representations.

Thu 24 SeptComputation and Language
The gist
The paper looks at how well large language models (LLMs) can turn formal meaning representations into English sentences and vice versa. The authors compare LLMs to a specialized system called ACE that parses sentences into meaning graphs and generates sentences from these graphs. They find that LLMs can generate English sentences from meaning representations surprisingly well, even better than some trained systems. However, LLMs have a much harder time parsing English sentences back into these formal meaning graphs, with much lower accuracy. This means that while LLMs can produce fluent sentences from semantic info, they do not necessarily understand the deeper semantic structure fully.
Open → 2609.30071v1

Sparse autoencoders reveal parts of speech in language data

Parts-of-Speech as Emergent Categories in SAE Latent Space

Abstract: Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

Thu 24 SeptComputation and Language
The gist
The paper studies how a type of AI model called a Sparse AutoEncoder (SAE) represents parts of speech like verbs and nouns. The authors find that these models can identify parts of speech well, but not by matching one part-of-speech category to one specific signal inside the model. Instead, groups of signals together represent these categories. The system organizes language information in a complex, spread-out way rather than using simple, single features.
Open → 2609.29362v1

Large language models show limits in mimicking infant syntax learning

A retrospective analysis on the use of LLMs to study infant syntax learning

Abstract: Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.

Tue 22 SeptComputation and Language
The gist
This paper looks back at how large language models (LLMs) have been used to study how babies learn sentence structure. The authors point out that the way these models are trained and tested involves a lot of assumptions that limit what we can conclude. They also find that using baby-like learning data doesn’t notably improve model performance on usual language tests. This suggests that actual infant learning processes might be quite different from how LLMs work.
Open → 2609.26539v1

Neural dynamics must respect language structure for real comprehension

Formal Properties of Language as Constraints on Neural Dynamics

Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural activity. Yet predictive success leaves mechanisms under-constrained. Here, we show that algebraic properties of language specify invariants that mechanisms must preserve: non-associative hierarchical grouping, commutativity, recursive closure, access to substructures, and structured workspace transitions. We term this the Neural Admissibility Program (NAP). Syntactic structure building is analyzed algebraically, with candidate mechanisms offered for each requirement: content-addressable workspace memory, graph-structured transient dynamics scheduling structure-building operations (e.g. stable heteroclinic channels), and a phase-coupled sealing operation recording grouping. Simulations show that a corrected Marcolli-Berwick entropy-optimized binding gate preserves grouping only within a narrow commitment band. As an alternative, we propose a novel neural binding operation we term 'Meld': two constituent populations converge through shared synapses, integrate sublinearly, and saturate. Meld is, to our knowledge, the closest neurally plausible composition law to syntactic Merge. It preserves every NAP invariant, uses known cortical operations, and recovers hierarchical structure at every tested depth and temperature. It predicts that effective population dimensionality separates alternative bracketings and that the composite depends on constituent disagreement. Importantly, the laws decoding bracketing most accurately are a priori inadmissible, showing that decoding accuracy alone cannot adjudicate between mechanisms. By specifying how neural dynamics can remain faithful to linguistic structure, the NAP changes the criterion by which neural implementations of cognition are evaluated.

Sun 13 SeptComputation and Language
The gist
Understanding language requires certain basic rules to be followed by brain processes, such as keeping words grouped properly and recognizing repeated patterns. The authors propose a formal framework called Neural Admissibility Program (NAP) that spells out these essential rules as mathematical properties. They test different neural mechanisms and introduce a new one called Meld, which closely matches how language structures work in the brain. Their work shows that just measuring how well a model predicts brain data is not enough to understand if it truly captures language processing.
Open → 2609.14384v1

LLM agents help analyze language structures faster and more widely

LLM Agents as Computational Typologists

Abstract: Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale crosslinguistic comparison labor-intensive and unscalable. We introduce AUTOTYPOLOGIST, an LLM agent for evidence-grounded typological analysis over reference grammars. The agent is capable of retrieving relevant grammar sections, analyzing interlinear glossed text (IGT), and iteratively reasoning over typological hypotheses using a ReAct-style workflow. We evaluate the system on TYPOLOGICAL FEATURE CODING against expert annotations and TYPOLOGICAL HYPOTHESIS TESTING with typological universals using 25 open-source reference grammars. Operating under different information constraints in TYPOLOGICAL FEATURE CODING, the agent can synthesize information from reference grammar prose but still faces challenges with only IGTs in the target language. In TYPOLOGICAL HYPOTHESIS TESTING, the agent can synthesize crosslinguistic evidence and identify both supporting cases and counterexamples. These findings suggest that LLM agents can support scalable and inspectable typological analysis, while still requiring expert validation.

Mon 7 SeptComputation and Language
The gist
Studying how languages differ is hard and takes a lot of expert work. The authors created AUTOTYPOLOGIST, a language model agent that can read grammar books and examples to find patterns in many languages. It checks facts carefully and uses step-by-step reasoning to understand complex language features. While it does well when it has full grammar texts, it struggles with just raw language examples and still needs experts to double-check its findings.
Open → 2609.07791v1