LLM reasoning accuracy varies across sampled geometry problem answers
Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence
Summary
When large language models (LLMs) try to solve formal geometry problems, taking multiple guesses can find the right number but not always a trustworthy explanation or confirmed solution. The authors introduce a way to measure three things: how often the right answer appears (coverage), how correctly the model reads those answers (realization), and how well an expert judge thinks the reasoning supports the answer (validity evidence). They tested this on a hard set of 441 problems and found that simply having the right answer in samples doesn’t guarantee accurate or well-supported solutions. These three measures capture different aspects of quality and should be reported separately.
What this means in practice
- •For llm developers: Measure and report separate scores for answer availability, reading accuracy, and reasoning support to improve sampled reasoning evaluation.
- •For automated theorem proving teams: Enhance solver evaluation by distinguishing between correct answer presence and the quality of derivation support during model reasoning.