Python dependency failures cut by replaying known package setups

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

Artificial IntelligenceMultiagent SystemsSoftware Engineering

Summary

When Python programs fail because their required packages don't work together, it can be tricky to fix the problem. The authors created a system called PLLM+ that first tries easy methods like reusing past package setups that worked before and checking package versions carefully. Only if those don't work does it try a more complex process using AI language models to suggest fixes. This approach solved more than half of the test cases and was much faster than just using the AI model alone. Most fixes came from reusing old working package setups, showing that simple reuse works well before trying AI.

What this means in practice

  • For python developers: Automatically fix dependency failures in Python code by reusing past successful package combinations with fallback AI repair.
  • For python package maintainers: Check and validate package version constraints by replaying known good configurations to prevent compatibility issues.

Authors

Veronica Poweska, Ariana Oyanguren, Jessica Pourleyli, Sourena Khanzadeh, Manar Alalfi

Abstract

Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 dependency-failing snippets. PLLM+ prioritizes inexpensive deterministic steps before invoking LLM-based repair: static AST-based interpreter inference, replay of historically successful dependency configurations from the competition-provided solutions database, and live PyPI validation of candidate package versions. When these steps do not resolve a case, the system falls back to a structured LLM-based repair loop with typed error classification and Proposer/Critic agents. On HG2.9K, PLLM+ solves 1,500 out of 2,891 snippets, compared with 1,169 solved by the PLLM baseline. It also reduces average runtime from 368.7 to 71.8 seconds per snippet. Most successful fixes come from replaying known configurations: 1,495 of the 1,500 successful fixes are produced by the solutions database, while the LLM fallback accounts for 5 additional fixes. These results suggest that, in this benchmark setting, deterministic reuse of previously validated dependency configurations is a simple and effective strategy, with LLM-based repair serving as a secondary fallback for cases not covered by prior solutions.