Benchmark reveals gaps in test generation across programming languages

Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation

Software Engineering

Summary

This paper looks at how tools can automatically create tests from bug reports in many programming languages. The authors made a large benchmark with nearly 2,000 examples in eight languages like Python, Java, and C++. They found that current AI methods work best for Python and struggle the most with C++. Their study highlights specific challenges both unique to certain languages and common across all languages when creating these tests. This work helps improve tools that verify software fixes across many programming environments.

What this means in practice

Authors

Kazuki Kusama, Sota Nakashima, Haruka Tokumasu, Masanari Kondo, Lingming Zhang, Yasutaka Kamei

Abstract

Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce MULTI-SWT-BENCH, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Using this benchmark, we conduct an empirical study of state-of-the-art LLMs with four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and perform a failure analysis across programming languages. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the success rate on Python exceeds the aggregate success rate across all languages, while C++ exhibits particularly low success rates. Our failure analysis identifies both language-specific challenges arising from repository testing conventions and cross-language challenges in inferring implicit setup requirements and preserving the target behavior through iterative revisions. These findings demonstrate the importance of multilingual evaluation and provide actionable directions for developing reproduction test generation methods that generalize across software ecosystems and reliably capture issue-specific behavior.