Papers for

puzzle app developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AI and humans choose different valid answers on reasoning puzzles

Investigating Human--AI Discrepancies via Multiple-Solution Problems

Abstract: Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and model reasoning lead to different distributions over valid solutions. Our testbed comprises 270 reasoning puzzles across five puzzle families. These multiple-solution puzzles each have 3 to 8 valid solutions and are simple enough that humans and models can solve them reliably. The resulting distributions differ markedly: models differ from one another, yet resemble each other far more than they resemble humans. Model distributions are, moreover, within every puzzle family, less diverse than human ones. We compare these discrepancies across puzzle categories, and trace how they respond to reasoning-effort settings, to prompting, and to perturbations of the puzzle that leave its solutions unchanged. Together, these results point at significant differences between human and AI problem-solving processes, and their choice among equally defensible solutions. As progressive deployment of AI systems in society comes into focus, evaluating such differences (beyond one-dimensional accuracy metrics) is increasingly important. Data and code are available at https://hai-discrepancies.github.io/

Mon 28 SeptComputers and SocietyArtificial Intelligence
The gist
Many problems can have several correct answers, but AI and humans tend to pick different ones. The authors studied 270 puzzles that each have multiple valid solutions, comparing how AI models and humans distributed their answers. They found AI models tend to be less diverse and more similar to each other than to humans. This shows AI and human problem-solving favor different solutions even when many are correct.
Open → 2609.34258v1