Agentic AI generates reliable Linux utility programs with human oversight

A Study of the Reliability of Agentic AI-Generated Programs

Software EngineeringArtificial Intelligence

Summary

This study looked at whether programs written by AI can be trusted as much as those written by people. The researchers asked AI tools to create ten well-known Linux utility programs and then tested these AI-made programs with special software tests. They found that the AI-generated code was usually just as reliable or even better than human code, having fewer memory problems but being more prone to getting stuck in loops. However, the quality of AI code depends a lot on how skilled the human guides the AI and what instructions they give. The study shows AI can help make solid software if humans carefully watch and guide the process.

What this means in practice

  • For software development teams: Use agentic AI tools under human supervision to produce reliable Linux utility programs with fewer memory errors compared to traditional code.
  • For software reliability engineers: Employ fuzz testing combined with AI-generated code to identify different class failures, such as infinite loops, improving testing coverage.

Authors

Ayesha Shafique, Barton P. MIller, Elisa R. Heymann

Abstract

Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a concrete ground truth. Third, we based our measure of reliability on a widely used testing technique, fuzz random testing. For this testing, we used both classic black box, generational testing and more modern coverage guided (gray box, mutational) testing using AFL++. We found that the AI-generated versions of the utility programs were typically as reliable - often more reliable - than the latest human-generated versions of these programs. While the AI-generated versions did have some failures, they were less common than the code from the standard repositories. Interestingly, the AI-generated code was less likely to have failures such as memory errors (such as buffer overflows) but more likely to have hangs such as infinite loops. In addition, we verified that generating robust and reliable software using agentic AI requires careful practice and human supervision. The quality of the code is highly dependent on the prompts and skills used, and how the human directing the process responds. We also demonstrated that using agentic AI workflow for software development (with its prompts and skills) can become a specification of the code that leads to cost-effective sustainability of the software.