Papers for

marketing analytics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AI moderation matches humans and finds more customer needs

AI-Moderated Interviews for Market Research and Digital Twins Calibration

Abstract: AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer "digital twins." Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).

Thu 24 SeptComputers and SocietyArtificial IntelligenceHuman-Computer Interaction
The gist
Finding out what customers want is important but expensive when done by humans. The authors tested AI-led interviews against human and simple surveys with real participants. They found that AI interviews gathered as many insights as human ones and even discovered more customer needs when the budget was the same. However, people felt more emotionally connected talking to humans. The AI data helped make digital models of customers, which predicted responses better than simple demographic profiles but not better than static surveys.
Open → 2609.29143v1

Large language model improves measuring consumer values from shopping behavior

Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

Abstract: Human values are deep motivational orientations that shape human behaviors. In e-commerce, they reveal the stable drivers behind users' purchase decisions. Compared with short-term interests, consumer values better explain how users evaluate products before purchase. However, consumer values are often implicit in complex and fragmented behavioral trajectories, leaving value measurement from e-commerce behaviors largely underexplored. To this end, we propose the Behavior-to-Value (B2V) task, which aims to identify consumer values from e-commerce behavioral trajectories. Centered on this task, we first construct the E-commerce Consumption Value Taxonomy (ECVT) and introduce B2V-Bench, the first B2V dataset and benchmark, based on anonymized Taobao behavioral logs. B2V-Bench consists of real-world purchase decision episodes, covering 25 types of purchase behaviors, along with corresponding consumer value orientations manifested in each episode. To improve consumer value measurement accuracy, we further present B2V-Verifier, a behavior-to-value measurement model based on Value Verification Tuning, which learns to assess whether behaviors provide sufficient evidence for each value inference. Experiments show that B2V-Verifier outperforms strong LLM baselines, improving multi-label classification by 34\%. The dataset and code will be publicly released upon acceptance.

Wed 16 SeptComputation and LanguageMachine Learning
The gist
People’s values influence what they buy, but these values are hidden in the complicated paths of their online shopping actions. The authors created a way to identify these values from shopping behavior using a new dataset based on Taobao logs. They developed a model, called B2V-Verifier, that better detects consumer values from these behaviors than previous language models, improving accuracy significantly. This work helps understand why consumers make certain purchase decisions.
Open → 2609.18203v1

AI generated data can change causal experiment questions and results

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

Abstract: AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.

Tue 15 SeptArtificial Intelligence
The gist
When artificial intelligence creates data from things like notes or images, it can change what a study is actually measuring, especially in experiments done step-by-step. The authors explain that these AI-created features can play many different roles in a study, and confusing these roles can lead to mistakes in interpreting results. They propose a clear way to classify AI-generated data and keep track of what question the study answers. Their approach helps avoid errors that come from mixing up treatment effects with other types of data changes. They also tested their ideas with simulations showing their methods can improve analysis when the AI data keeps important information.
Open → 2609.17772v1

Interpretable random forests improve treatment effect predictions and insights

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

Abstract: In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individuals than for others. To address this two-fold challenge of prediction and interpretation, we introduce an algorithm based on decision trees and random forests for estimating individual treatment effects. Our algorithm is simple: it operates exactly like a standard random forest, but with a different splitting criterion, and requires no additional workarounds such as double machine learning or orthogonalization as used in Generalized random forests. It handles observational studies with varying treatment propensities without requiring separate estimation of the full propensity function. This is achieved by combining two splitting criteria---one targeting heterogeneity in the treatment effect, the other targeting bias correction for the average treatment effect---which together improve split point selection and automatically distinguish confounders from features responsible for heterogeneity. As a result, interpretation follows directly from the fitted tree structure itself, that is, from which features the trees split on and with which split statistics, without requiring separate post-hoc analysis. For the theoretical analysis of this algorithm, we consider a change point model with step functions for potential outcomes and treatment propensity and provide insights into the theoretical underpinnings of our approach. Simulation studies show that our simple algorithm achieves comparable, and often better, prediction accuracy than existing methods, while substantially improving interpretability.

Tue 15 SeptMachine Learning
The gist
Predicting who benefits most from a treatment is important but hard to understand. The authors created a simple tree-based method that not only predicts effects well but also explains which factors cause differences between people. Their approach works with biased data without extra complicated steps and shows results directly through the tree structure. Tests show it predicts as well or better than existing methods while being easier to interpret.
Open → 2609.16971v1

Large language models improve social network modeling and dynamics

LLMs for Social Network Modeling: From Network Generation to Dynamic Processes

Abstract: Large language models (LLMs) are rapidly emerging as a new paradigm for modeling social networks by representing users and their relationships and interactions through natural language. Unlike classical network models or deep learning approaches, LLMs can simulate context-aware social behavior and language-driven interactions, enabling more realistic modeling of network formation and dynamic social processes. However, existing studies are scattered across different research communities and lack a unified perspective. This survey presents the first comprehensive review of LLMs for social network modeling by organizing the literature into two broad categories: network generative models and dynamic process models. Network generative models are further classified into selection-based and interaction-based approaches, while dynamic process models are categorized into opinion dynamics, information diffusion, and rumor propagation, each with their underlying modeling mechanisms. LLMs enable rich textual social interactions and decision-making, but they also exhibit many limitations, including inherent social biases and prompt sensitivity. We outline these open research challenges and discuss future directions in LLM-based social network modeling.

Mon 7 SeptSocial and Information NetworksArtificial Intelligence
The gist
Social networks are complicated, with people interacting and sharing opinions in many ways. The authors survey how large language models (LLMs) can represent users, their relationships, and conversations using natural language. These models help simulate how social networks form and evolve, including how information and rumors spread or how opinions change. The paper groups current research into categories based on network creation and dynamic social behaviors, while noting challenges like bias and sensitivity to input. They also outline directions for improving the use of LLMs in social network studies.
Open → 2609.08049v1