How Do Software Professionals Evaluate AI-Generated Code? (Registered Report)

2026-07-10Software Engineering

Software Engineering
AI summary

The authors study how software developers judge the quality of code created by AI tools like GitHub Copilot and ChatGPT. They use surveys and interviews to gather detailed information from professionals about their experiences and opinions. Their goal is to develop a theory based on real user views about how people decide if AI-generated code is good or not. This helps understand what factors influence trust and acceptance of AI in coding.

Generative AIGitHub CopilotChatGPTCode evaluationConstructivist grounded theorySemi-structured interviewsSurvey methodologyTheoretical saturationSoftware development practices
Authors
Samuli Määttä, Hera Arif, Burak Turhan, Paul Ralph, Markus Kelanti
Abstract
Recent advances in generative AI tools have significantly changed how software professionals write, evaluate, and interact with code. Generative AI tools such as GitHub Copilot, ChatGPT, and Claude are increasingly being integrated into everyday workflows. Despite the growing adoption of and reliance on these tools, it remains unclear as to how software professionals evaluate the code they generate. To explore this topic, we will conduct a constructivist grounded theory study that incorporates a survey, semi-structured interviews, and laddering interviews. With the initial survey data collection complete, we aim to interview 20--50 software professionals iteratively until theoretical saturation is achieved. This research aims to build a theory of how software professionals evaluate AI-generated code, grounded in their accounts of evaluative practices, perceptions, and preferences.