Machine learning finds swapped arguments in Python code without full context

Detecting Argument-Swap Bugs Using Context-Enhanced Code Representations

Software Engineering

Summary

Sometimes programmers accidentally swap the order of inputs when calling functions, which can cause bugs. The authors created a method called BugProbe that uses machine learning to spot these swapped arguments in Python code, even when it's hard to find out which function is being called. They trained BugProbe with lots of example code and tested it with real bugs found in popular projects, showing it works better than earlier methods. This approach helps find bugs in dynamically typed languages like Python, where traditional methods struggle.

What this means in practice

  • For python developers: Detect swapped argument bugs during code review even when function definitions are unavailable or dynamic typing complicates analysis.
  • For software quality teams: Integrate BugProbe into testing pipelines to catch argument order mistakes that typical static analysis might miss in Python projects.

Authors

Subrata Das, Ali Aman, Muhammad Asaduzzaman, Kawser Wazed Nafi, Salimur Choudhury

Abstract

Names of source code elements convey rich semantic information and have been widely used in software engineering tasks such as bug detection, code completion, type prediction, and code classification. Prior studies exploit lexical similarity between method arguments and formal parameter names to detect bugs caused by incorrectly ordered arguments, typically relying on establishing mappings between method calls and their corresponding definitions. However, such mappings are often difficult to obtain in dynamically typed languages like Python. In this paper, we present BugProbe, a learning-based approach for detecting incorrectly ordered arguments in Python method calls that does not require call-to-definition mappings. Our approach leverages multiple sources of contextual information, including local context and argument usage context, and combines name-based similarity with machine learning to construct expressive representations of method arguments. We collect a new dataset of 132,739 Python source files from the top-1,000 starred GitHub repositories, yielding 3,371,244 synthetic training examples, and contribute a curated benchmark of 55 real-world argument-swap bugs manually verified from commit histories. We evaluate our approach on this dataset and show that it achieves high accuracy and consistently outperforms a state-of-the-art baseline across standard evaluation metrics. These results demonstrate that effective detection of argument-ordering bugs is possible without relying on explicit call-to-definition resolution, making the approach well suited for dynamically typed language settings.