Mining real incomplete software changes to improve co-change detection

The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects

Software Engineering

Summary

Software changes often need to be made in many files together, but developers can miss some of these related changes. Previous methods tried to find rules about which files usually change together, but tested on fake incomplete changes instead of real ones. The authors created a new dataset of real incomplete software changes by looking at bugs and their fixes from open-source projects. They used this data to study how common missed changes are and improved testing for an existing rule-finder method.

What this means in practice

Authors

Savira Ramadhanty, Profir-Petru Pârţachi, Yoshiya Ishida, Takashi Kobayashi

Abstract

During software development, a modification to a software component may propagate across the system, requiring precise identification and correct revision of all affected components. This is a complex task, and developers often (45.7%) miss related changes. To address this, several methods have been developed to extract co-change rules from files that are frequently changed together in the revision history. However, previous research evaluated the methods using artificially created incomplete changes, which may not be representative of real-world data. To solve this problem, we construct a dataset by mining incomplete changes from a collection of open-source software, using information about induced bugs and their respective fixes from an issue tracking platform. We also analyze the characteristics of incomplete changes using this constructed dataset and found that 89.4% of missed changes involved five or fewer files. Finally, we re-evaluate LCExtractor, an existing co-change rule extraction method, on our constructed dataset, and we identify the optimal sorting criterion and the impact of the number of used commits.