Papers for

python developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

A verified stream protocol combines sync and async modules smoothly

Designing a Producer-driven Stream Protocol by Formal Refinement

Abstract: The coroutine has broadly diffused in the practice of concurrent programming in the form of generators and asynchronous functions as well as processes communicating through pipes. We wanted to use coroutines in Python to create single-threaded Unix-style pipelines. Unfortunately, available solutions in Python are cumbersome to use. The JavaScript push-stream protocol appeared to be a good alternative. However, its specification is incomplete and ambiguous. We have used TLA$^{+}$ and the TLC model checker to re-derive the protocol and obtain protocol-specific verification tools. In this paper, we present a formal specification of a push-stream protocol that 1) seamlessly combines synchronous and asynchronous modules, encapsulating the choice within each module; 2) provides flow control without using bounded buffers; 3) gracefully and unambiguously terminates; 4) does not require dynamic allocation of objects on the heap. In addition to completely describing expected behaviours, our specification improves on the original design by 1) allowing the input and output of intermediate pipeline modules to terminate independently and 2) explicitly reporting when a module is pending on the execution environment, to avoid incorrect resuming. We specify the protocol as a sequence of refinement steps and derive by equivalence a specification of what an abstract module may do. We then refine the latter into a module checker that can verify concrete module specifications for conformity. We have verified the key properties of all specifications and the validity of refinement steps with TLC. In supplemental material, we provide all TLA$^{+}$ specifications, show that the protocol is sufficiently expressive to implement a superset of all original JavaScript modules, as well as a performance comparison with Python alternatives and Unix pipes.

Sun 27 SeptProgramming Languages
The gist
Programming with pipelines that mix different ways of running tasks can be tricky in Python because current tools are hard to use. The authors studied an existing JavaScript stream protocol that uses a pushing approach and found it unclear. They carefully defined and checked a clearer version using formal math-based methods and software tools, improving how parts of the pipeline start and stop independently and handle waiting. Their work helps ensure that stream pipelines behave correctly and efficiently without needing complex memory use.
Open → 2609.33813v1

Graph guided method improves software environment setup success

Graph-Guided Repository Environment Construction

Abstract: Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifacts and may only become apparent during execution. Existing agent-based approaches address this problem through iterative interaction, but information about the current construction state, including discovered requirements, satisfied and unresolved prerequisites, and their dependencies, can remain distributed across the interaction history. We present Graph2Env, an agent-based approach centered on DepGraph, a typed dependency graph that explicitly represents the environment requirements needed for repository execution, their dependency relations, and their states. Graph2Env uses DepGraph to guide environment construction and continuously refines it with execution feedback, while persisting successful repairs into a replayable construction procedure. The resulting artifacts are then applied in a fresh environment to verify that the constructed environment can be reproduced. We evaluate Graph2Env on a benchmark of 200 Python repositories drawn from RATBench and EnvBench, against a static dependency-inference baseline (pipreqs), three specialized environment-construction systems (Repo2Run, RAT, and SetupX), and two general-purpose coding agents (SWE-agent and Claude Code). Graph2Env achieves an 81.0% Environment Build Success Rate (EBSR) and a 59.3% Environment Setup Success Rate (ESSR), outperforming the strongest baseline by 9.5 and 9.0 percentage points, respectively.

Sun 27 SeptSoftware EngineeringArtificial Intelligence
The gist
Getting software project environments ready to run code properly is hard because requirements are spread across different files and may only show up when testing the code. The authors created Graph2Env, a tool that maps out all the needed parts and their connections in a clear graph to help build these environments step-by-step. By learning from errors during execution and keeping track of what works, Graph2Env builds setups that can be repeated reliably. When tested on many Python projects, it did better than existing methods at preparing working environments.
Open → 2609.33429v1

Python dependency failures cut by replaying known package setups

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

Abstract: Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 dependency-failing snippets. PLLM+ prioritizes inexpensive deterministic steps before invoking LLM-based repair: static AST-based interpreter inference, replay of historically successful dependency configurations from the competition-provided solutions database, and live PyPI validation of candidate package versions. When these steps do not resolve a case, the system falls back to a structured LLM-based repair loop with typed error classification and Proposer/Critic agents. On HG2.9K, PLLM+ solves 1,500 out of 2,891 snippets, compared with 1,169 solved by the PLLM baseline. It also reduces average runtime from 368.7 to 71.8 seconds per snippet. Most successful fixes come from replaying known configurations: 1,495 of the 1,500 successful fixes are produced by the solutions database, while the LLM fallback accounts for 5 additional fixes. These results suggest that, in this benchmark setting, deterministic reuse of previously validated dependency configurations is a simple and effective strategy, with LLM-based repair serving as a secondary fallback for cases not covered by prior solutions.

Tue 22 SeptArtificial IntelligenceMultiagent SystemsSoftware Engineering
The gist
When Python programs fail because their required packages don't work together, it can be tricky to fix the problem. The authors created a system called PLLM+ that first tries easy methods like reusing past package setups that worked before and checking package versions carefully. Only if those don't work does it try a more complex process using AI language models to suggest fixes. This approach solved more than half of the test cases and was much faster than just using the AI model alone. Most fixes came from reusing old working package setups, showing that simple reuse works well before trying AI.
Open → 2609.26952v1

Python projects face common issues running on different operating systems

An Empirical Analysis of Cross-OS Portability Issues in Python Projects

Abstract: While Python is designed as a cross-platform language, real-world applications encounter portability failures when deployed across different operating systems. We present the first large-scale empirical study of cross-OS portability issues in Python, analyzing 2,042 open-source repositories using two complementary approaches: systematic cross-OS test reexecution and manual analysis of GitHub issues. Our cross-platform testing of 500 projects reveals that 11.2% exhibit OS-dependent test failures. Through systematic analysis of 240 GitHub issues, we confirm 102 genuine portability problems spanning 95 additional projects. We develop a comprehensive taxonomy identifying 7 primary failure categories - with file/directory operations, process management, and library dependencies being most prevalent - along with 24 distinct sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns. Our evaluation reveals that existing static analysis tools provide minimal support for portability detection, while large language models achieve 40-79% accuracy in identifying issues and 50-77% success in generating fixes when provided with structured guidance. Through 33 contributed pull requests, we demonstrate practical applicability and developer acceptance (17 merged, zero rejected) of our findings. This work establishes the first comprehensive baseline for understanding and addressing cross-OS portability issues in Python, providing actionable insights for developers, tool designers, and the broader research community.

Tue 22 SeptSoftware Engineering
The gist
Python is meant to work the same on any computer system, but many real programs still run into problems when moved between different operating systems like Windows and Linux. The authors studied over two thousand projects and found about 11% had test failures due to OS differences. They categorized the main types of problems and tested how well current tools and AI can detect and fix these issues. Their findings include useful patterns for developers and evidence that some fixes are accepted by the open-source community.
Open → 2609.25531v1

Machine learning finds swapped arguments in Python code without full context

Detecting Argument-Swap Bugs Using Context-Enhanced Code Representations

Abstract: Names of source code elements convey rich semantic information and have been widely used in software engineering tasks such as bug detection, code completion, type prediction, and code classification. Prior studies exploit lexical similarity between method arguments and formal parameter names to detect bugs caused by incorrectly ordered arguments, typically relying on establishing mappings between method calls and their corresponding definitions. However, such mappings are often difficult to obtain in dynamically typed languages like Python. In this paper, we present BugProbe, a learning-based approach for detecting incorrectly ordered arguments in Python method calls that does not require call-to-definition mappings. Our approach leverages multiple sources of contextual information, including local context and argument usage context, and combines name-based similarity with machine learning to construct expressive representations of method arguments. We collect a new dataset of 132,739 Python source files from the top-1,000 starred GitHub repositories, yielding 3,371,244 synthetic training examples, and contribute a curated benchmark of 55 real-world argument-swap bugs manually verified from commit histories. We evaluate our approach on this dataset and show that it achieves high accuracy and consistently outperforms a state-of-the-art baseline across standard evaluation metrics. These results demonstrate that effective detection of argument-ordering bugs is possible without relying on explicit call-to-definition resolution, making the approach well suited for dynamically typed language settings.

Tue 15 SeptSoftware Engineering
The gist
Sometimes programmers accidentally swap the order of inputs when calling functions, which can cause bugs. The authors created a method called BugProbe that uses machine learning to spot these swapped arguments in Python code, even when it's hard to find out which function is being called. They trained BugProbe with lots of example code and tested it with real bugs found in popular projects, showing it works better than earlier methods. This approach helps find bugs in dynamically typed languages like Python, where traditional methods struggle.
Open → 2609.17844v1