Papers for

business data analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Data agents with workflows improve reliability in AI data science

Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

Abstract: Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.

Mon 28 SeptArtificial Intelligence
The gist
Doing reliable data analysis using AI is harder than just having a smart program. The authors show that breaking down AI data work into clear steps like seeing the data, planning, doing the work, checking results, and fixing problems helps make AI data scientists more dependable. They identified common problems that cause quiet errors and suggested better ways for AI to keep track of its knowledge and fix errors. They organized what is known about data agents and shared useful tools for others to build on.
Open → 2609.35255v1

Control-based forecasting improves accuracy for changing time series

CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning

Abstract: Time series forecasting underpins critical decision-making across diverse domains. While large language models (LLMs) offer promising reasoning capabilities, existing LLM-based time series forecasting approaches either reduce them to numerical predictors that bypass their strengths, or allow direct forecast generation that destabilizes predictions in non-stationary settings. We introduce CTRL, a framework that decouples semantic reasoning from quantitative prediction. A frozen backbone generates base forecasts, while specialized LLM agents function as controllers that analyze backbone prediction errors through decomposed trend, seasonal, and irregular components, grounding reasoning in interpretable temporal structure. Each agent outputs compact control signals that a lightweight residual decoder translates into forecast corrections. CTRL incorporates label-free test-time adaptation that detects distribution shift from input statistics alone and readapts control signals with only 3-24 LLM calls via caching. CTRL is explicitly designed to improve robustness under non-stationary temporal dynamics and distribution shift, while remaining competitive on highly stationary time series where adaptive correction provides limited additional benefit.

Sun 20 SeptMachine LearningComputation and Language
The gist
Time series forecasting involves predicting future data points based on past observations, which is important in many areas like weather or finance. The authors show that large language models (LLMs), which are good at reasoning, have so far been used only as simple predictors or direct forecasters that struggle when conditions change. They introduce CTRL, which separates reasoning from prediction by letting a fixed forecast model make a base guess and then using LLMs as controllers to understand and adjust errors in the forecast’s trend, seasonality, and irregular parts. This method helps make predictions more stable and accurate when the data’s behavior shifts over time.
Open → 2609.23257v1

AI coders debate and agree to improve qualitative coding accuracy

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

Abstract: The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.

Thu 10 SeptHuman-Computer InteractionArtificial Intelligence
The gist
Figuring out how AI can help people label text data is tricky because it depends on many things. The authors studied how multiple AI agents code, argue, and come to agreements when labeling text from different sources. They found that longer instructions, how similar the data is, and how much the AI agents disagree affect how well the AI does. Surprisingly, harder debates between AI agents sometimes meant the final answers were more accurate. The authors also noticed AI behaves a bit like humans in discussions but doesn't adapt well to changing contexts.
Open → 2609.11109v1