Data agents with workflows improve reliability in AI data science
Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
Artificial Intelligence
Summary
Doing reliable data analysis using AI is harder than just having a smart program. The authors show that breaking down AI data work into clear steps like seeing the data, planning, doing the work, checking results, and fixing problems helps make AI data scientists more dependable. They identified common problems that cause quiet errors and suggested better ways for AI to keep track of its knowledge and fix errors. They organized what is known about data agents and shared useful tools for others to build on.
What this means in practice
- •For software engineering teams: Build AI-powered data tools with structured workflows to improve error detection and recovery in automated analysis.
- •For business data analysts: Use AI agents that manage data workflows to help ensure trustworthy data insights and reduce silent failures.
A survey. It maps existing work.
Authors
Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai, Zijin Hong, Chuang Zhou, Zheng Yuan, Qinggang Zhang, Qing Li, Xiao Huang
Abstract
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.