Abstract: Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce MULTI-SWT-BENCH, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Using this benchmark, we conduct an empirical study of state-of-the-art LLMs with four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and perform a failure analysis across programming languages. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the success rate on Python exceeds the aggregate success rate across all languages, while C++ exhibits particularly low success rates. Our failure analysis identifies both language-specific challenges arising from repository testing conventions and cross-language challenges in inferring implicit setup requirements and preserving the target behavior through iterative revisions. These findings demonstrate the importance of multilingual evaluation and provide actionable directions for developing reproduction test generation methods that generalize across software ecosystems and reliably capture issue-specific behavior.