LLM reasoning accuracy varies across sampled geometry problem answers

Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence

Artificial Intelligence

Summary

When large language models (LLMs) try to solve formal geometry problems, taking multiple guesses can find the right number but not always a trustworthy explanation or confirmed solution. The authors introduce a way to measure three things: how often the right answer appears (coverage), how correctly the model reads those answers (realization), and how well an expert judge thinks the reasoning supports the answer (validity evidence). They tested this on a hard set of 441 problems and found that simply having the right answer in samples doesn’t guarantee accurate or well-supported solutions. These three measures capture different aspects of quality and should be reported separately.

What this means in practice

  • For llm developers: Measure and report separate scores for answer availability, reading accuracy, and reasoning support to improve sampled reasoning evaluation.
  • For automated theorem proving teams: Enhance solver evaluation by distinguishing between correct answer presence and the quality of derivation support during model reasoning.

Authors

Xiao Yue, Guangzhi Qu

Abstract

Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Coverage is answer availability, realization is readout accuracy on the frozen candidate pool, and validity evidence is a label-blinded critic judgment of derivational support rather than a proof certificate. CRV freezes each candidate pool before comparing readouts and analyzes covered failures by correct-answer multiplicity and within-problem discrimination. On HardShift441, a 441-problem set for which a reference solver leaves 406 problems unsolved, a LoRA-adapted Qwen2.5-7B generator obtains 24.2% average single-sample accuracy and 68.9% pass@16, whereas verifier-weighted self-consistency (WSC) reaches 38.0%. Readout accuracy is particularly low when the correct answer occurs only once or twice in the pool. In a separate constructed audit of 195 covered problems, the critic labels 12 correct-answer representatives as supported, 181 as refuted, and two as uncertain. These results show that coverage, realization, and validity evidence from the critic are distinct quantities and should be reported separately.