Papers for

content management teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Knowledge pull requests improve continual document updates across languages

Knowledge Pull Requests for Continual Document Authoring

Abstract: We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content, producing a ChangeLog that separates what knowledge changes (claim proposal) from how the text changes (document diff). We evaluate KPRs on revising Wikipedia across languages and updating query-driven reports on RAGTIME. KPRs integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated. A KPR-revised article also grounds question answering better than a frontier model with search, which does not surface knowledge documented only in other languages.

Tue 22 SeptComputation and Language
The gist
Documents like Wikipedia articles need constant updates as new facts appear. The authors present a method called Knowledge Pull Requests (KPRs) that helps update documents by clearly showing what knowledge changes and how the text is altered. KPRs work by extracting new claims, deciding where they fit in the document, and marking contradictions with old content. This approach preserves more existing information and adds more useful facts than rewriting entire texts from scratch. It also helps answer questions better by including knowledge from documents in different languages.
Open → 2609.26634v1

Closed loop system greatly improves meeting text output targets

A Closed-Loop Control Architecture for Reliable Constraint Satisfaction in LLM Text Generation

Abstract: Software systems increasingly embed a large language model in features that must satisfy a numeric output constraint, that is, a requirement expressible as a number or an interval and checkable by code, such as a target word count or a target readability grade band. Because such a model is non-deterministic, is configured through natural-language instructions rather than a typed interface, and satisfies a stated requirement only approximately, a single prompt neither reliably meets the target nor preserves the source content. This paper presents and evaluates a closed-loop control architecture for this problem. It has five stages: generate, evaluate, adjust, archive, and analyze. The model is called only to write and to edit text, while deterministic code compares a composite readability value against a target band, rejects any edit that drops source entities, numbers, or keywords, and makes every accept decision. Over 114 single-shot generation jobs and 240 closed-loop runs on four commercial models, single-shot prompting met the target in 21.1 to 31.6 percent of cases and the closed loop in 92.5 to 98.8 percent, within two edit rounds on average and at a recall-based fidelity of 0.92 to 0.93; the two models common to both settings show the same effect. Because the controller optimizes the value on which success is scored, the result establishes reproducible control over a declared, computable metric and not validated human difficulty. The transferable practice is to declare the acceptance condition as code, bound the model to local edits, and gate every edit on a content check.

Thu 17 SeptSoftware Engineering
The gist
When using large language models to create text that meets specific numeric rules—like word counts or readability scores—one attempt often doesn’t get it right. The authors show a system that generates text, checks if it fits the rules, then edits it repeatedly until it does. This looping method made the output meet targets much more reliably and kept important details intact. The system works by using code to decide if edits are good enough, rather than relying on the language model’s guesswork alone.
Open → 2609.19710v1

Region aware retrieval improves image and text matching in large models

RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models

Abstract: Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20\% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.

Tue 15 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceInformation Retrieval
The gist
Finding specific parts of an image and matching them to related words or other image parts is important but hard. The authors designed RegRet, which helps large multimodal models look more closely at image regions without losing the overall picture. They also created a big new dataset to train and test this regional matching. Their experiments show RegRet does a better job than previous methods, especially when it comes to detailed image parts.
Open → 2609.16847v1

Text summarization improves by focusing on important tokens

TIAO: Token Importance-Aware Policy Optimization for Text Summarization

Abstract: Text summarization requires models to condense content while preserving key qualities such as consistency and coherence. Large language models (LLMs) have shown strong performance on this task and can be further improved through reinforcement learning (RL). However, most existing methods apply reward signals directly to undifferentiated token sequences, overlooking the varying importance of individual tokens to word and sentence level quality in summarization. In this paper, we propose Token Importance-Aware Policy Optimization (TIAO), a novel reinforcement learning strategy that explicitly leverages token-importance awareness. Specifically, TIAO identifies core tokens based on token dependency and reweights a trajectory's advantage according to its overall dependencies. Experiments on the real world dataset show that our TIAO achieves highly competitive results, and that a 7B foundation model enhanced by TIAO performs comparably to GPT-4 and GPT-5-nano. Code is available at https://github.com/TechCloud-x/TIAO

Tue 15 SeptComputation and Language
The gist
Summarizing text means making it shorter but still keeping the important ideas clear and connected. The paper introduces a new method called TIAO that helps computer models recognize which words in a summary are most important. By paying special attention to these key words during training, the method improves how well summaries make sense and capture the main points. Experiments show that even smaller trained models using this method can perform as well as much larger AI systems like GPT-4. This approach refines how AI learns to summarize by understanding the role of each token in the sentence.
Open → 2609.16748v1

CausalChapter improves chaptering for long instructional videos

CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

Abstract: Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from textualized video inputs, but they remain costly and brittle for content-dense lecture videos with long transcripts, smooth topic transitions, and detailed chapter outputs. A scalable segment-then-caption paradigm reduces this cost, but introduces two new challenges: boundary error propagation and fragmented cross-chapter context. We propose \textbf{CausalChapter}, an intervention-inspired framework for long-video chaptering that estimates prediction-level influence through lightweight masking and removal interventions. For boundary localization, our Local Dependency Shift module detects drops in predictive dependency between adjacent temporal windows; for chapter description generation, our Cross-Segment Support Selection module reranks historical contexts according to their support for the current prediction. Experiments on long-video chaptering benchmarks show that CausalChapter improves boundary localization, chapter description quality, and cross-chapter coherence.

Tue 8 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Long instructional videos can be hard to navigate because they often don’t have clear chapters to separate topics. The authors propose CausalChapter, a method that helps find chapter boundaries and write chapter summaries more accurately by checking how different parts of the video influence predictions. This approach uses a clever masking technique to spot where topics change and improves linking across chapters for better descriptions. Their method works better than previous ones for videos with long, dense content and smoothly changing topics.
Open → 2609.08686v1