Coding agents face new puzzle test requiring deep program reasoning

Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents

Software EngineeringArtificial Intelligence

Summary

Existing tests for coding AI mainly ask models to guess what a program will do, but these tests assume the AI can't run the program to check. The authors created Codoku, a puzzle game where AI must fill in missing parts of a program without being able to run it, forcing it to think through how the pieces fit together. Because puzzles are generated fresh each time and have guaranteed solutions, Codoku stays challenging and fair as AI gets better. Tests show even strong models struggle with the harder puzzles, showing this is a useful new way to measure coding reasoning skills.

What this means in practice

  • For ai developers: Test and improve AI coding agents' ability to reason about programs without running them using a renewable set of puzzles.
  • For software testing teams: Use Codoku to generate varied program reasoning challenges that detect subtle flaws in code analysis tools.

Authors

Cong Li, Hao Sun, Zenan Li, Zhendong Su

Abstract

Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input. Coding agents break two assumptions on which these benchmarks rest: an agent can recover the answer by executing the program instead of reasoning about it, and fixed task sets drawn from existing programs are increasingly exposed to contamination, yet costly to renew. We introduce Codoku (code sudoku), a renewable benchmark in which a solver fills typed cells in a partial program to satisfy global static and dynamic constraints, such as a prescribed control-flow graph and execution path. Because a partial program cannot be executed and valid fillings are sparse in an exponentially large space of interdependent choices, neither tool use nor enumeration can substitute for program reasoning. Puzzles are synthesized from scratch via semantic reification, so fresh puzzles of controllable complexity can be generated on demand, each with a witness that guarantees solvability. We evaluate five frontier models on 300 puzzles through a coding agent free to use any tool within a fixed budget. Small puzzles already challenge open-weight models, whereas even proprietary models solve only about half of the large ones. Codoku thus offers a renewable testbed for program reasoning that can keep pace with rapidly improving coding agents. GitHub: https://github.com/connglli/Codoku.