AI generated data can change causal experiment questions and results

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

Artificial Intelligence

Summary

When artificial intelligence creates data from things like notes or images, it can change what a study is actually measuring, especially in experiments done step-by-step. The authors explain that these AI-created features can play many different roles in a study, and confusing these roles can lead to mistakes in interpreting results. They propose a clear way to classify AI-generated data and keep track of what question the study answers. Their approach helps avoid errors that come from mixing up treatment effects with other types of data changes. They also tested their ideas with simulations showing their methods can improve analysis when the AI data keeps important information.

What this means in practice

  • For clinical trial designers: Improve experimental precision by correctly classifying AI-derived patient data roles to prevent bias in treatment effect estimates.
  • For marketing analytics teams: Avoid misinterpreting AI-generated customer behavior features by explicitly defining their causal roles in sequential campaign experiments.

Tested on simulated data.

Authors

Takes Fujita, Nobutaka Hattori

Abstract

AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.