More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs

2026-08-31Computation and Language

Computation and Language
AI summary

The authors studied how well large language models (LLMs) can tell if a math problem can be solved, focusing on different languages. They created a new benchmark that includes paired math problems in English, French, and Greek to test this ability. Their results showed that the models use a similar, language-independent way to judge solvability. Interestingly, even though English models solve math problems better, they are less faithful in signaling when a problem is solvable. This work helps understand how language affects math reasoning in AI.

Large Language ModelsMathematical ReasoningSolvability DetectionMultilingual BenchmarkSolvability BeliefFaithfulnessReliableMathLanguage-Agnostic FeaturesMathematical Problem Solving
Authors
Maria-Eleni Zoumpoulidi, Nikolaos Xiros, Georgios Paraskevopoulos
Abstract
Solvability detection is one of the most challenging aspects of mathematical reasoning for Large Language Models (LLMs). While prior work has studied this capability extensively, these analyses have been limited to English. Consequently, it remains unclear whether multilingual failures arise from differences in internal Solvability Belief or from language-dependent failures to express it. To address this gap, we introduce the first multilingual benchmark of paired solvable and unsolvable mathematical problems, extending ReliableMath to French and Greek. Using this, we train multilingual probes predicting Solvability Belief and analyze the solvability detection capabilities of state-of-the-art LLMs behaviorally, representationally, and in terms of faithfulness. We find that Solvability Belief is encoded as a largely universal, language-agnostic feature, and that higher-resource languages such as English, despite achieving stronger mathematical reasoning performance, exhibit lower solvability-detection faithfulness.