Reflection enhanced search improves code generation from natural language prompts
ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
Computation and Language
Summary
Writing code from descriptions can be tricky because early guesses might miss hidden details or keep making the same mistakes. The authors created a system called ReMCTS that searches smartly through different code tries, remembers past errors, and learns from them to avoid repeating problems. They tested this system on benchmarks where it did better than just generating code once, showing more success at making working programs. Their approach can also work with different programming languages and uses real test results to guide its decisions.
What this means in practice
- •For software development teams: Improve automatic code generation tools by integrating reflection-based search that remembers and learns from errors to produce more reliable code snippets.
- •For compiler developers: Incorporate execution feedback and debugging context in search strategies to enhance compiler-backed code synthesis and testing pipelines.
Authors
Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan
Abstract
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.