Papers for

qa system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Selective inference improves tree structured decision making in reinforcement learning

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Abstract: Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO

Mon 28 SeptArtificial Intelligence
The gist
When computers try to make decisions by exploring different possible future moves, they often pick choices based on past results. But this can cause a problem because some options look better just because they were selected before, not necessarily because they truly are better. The authors found a way to correct for this bias by carefully re-evaluating choices and their alternatives, making the decision-making process fairer and more accurate. Their method, called SIPO, showed better performance in question-answering tests using large language models.
Open → 2609.34805v1

Knowledge graph method measures large language model context understanding

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

Abstract: While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.

Thu 24 SeptArtificial IntelligenceMachine Learning
The gist
Large language models can produce impressive answers but it’s unclear if they truly understand the information around them or just guess based on patterns. The authors developed a way to check if these models really get the context by comparing their answers to knowledge graphs, which show facts and their connections. They created a special score to see how well the model's answers match the structure and meaning of these facts. This method also helps spot exactly where models make reasoning mistakes. Their tests on multiple benchmarks showed better results than previous tools.
Open → 2609.30484v1

Confidence signals in language models reveal instability in answers

When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models

Abstract: Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.

Tue 15 SeptMachine LearningArtificial Intelligence
The gist
This paper looks at two ways language models show how sure they are about their answers. One way checks how likely a model thinks its immediate answer token is, and the other counts how often the same answer shows up in many tries. The researchers found these two signals do not always agree and that the global confidence (the repeated-answer count) is better linked to correct answers. When the two signals differ a lot, it often means the model's answers are unstable or uncertain, especially on harder questions. This shows confidence is not one simple number but several measurements that tell different stories.
Open → 2609.16933v1