MechReasoner tests mechanistic reasoning with qualitative physics

MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics

Artificial Intelligence

Summary

Generating detailed explanations about how things work can be tricky because current AI models sometimes ignore the actual physical and structural rules involved. This paper presents MechReasoner, a tool that simulates physical mechanisms using qualitative reasoning and checks if AI answers follow the proper logic and constraints. The authors created a large set of test questions that grow more complex, showing even advanced AI models like GPT-5.5 struggle as the problems get harder. This means tools like MechReasoner can help build fair and understandable tests for mechanistic understanding in AI.

What this means in practice

  • For ai developers: Evaluate and improve AI models' understanding of physical mechanisms using a standardized simulator and benchmark with graded complexity.
  • For educational technology teams: Develop automated tools that assess mechanistic reasoning in students by referencing simulator-supported benchmarks for physics concepts.

Authors

Danilo Gusicuma, André Freitas

Abstract

This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.