Papers for

customer experience teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AI assistants rarely record buyer decisions after giving purchase advice

Purchase Advice and Observable Buyer Responses in Real AI Conversations

Abstract: How often does a generative assistant persuade someone to buy, or persuade them not to buy? Conversation logs contain recommendations, but they do not necessarily record subsequent decisions. We audit 317 historical interactions from Aiso's proprietary research database of licensed, consent-based, de-identified conversations with commercially available AI assistants. Single-agent AI-assisted screening identifies 68 purchase-directed records; collapsing one shared-prefix copy yields 67 retained episodes, dated April 2023 to July 2025. Assistant responses provide candidate options, acquisition channels, or conditional preferences in 52 episodes (77.6%). One episode contains conditional redirection away from a named accommodation candidate. No episode is coded as advice to abandon or defer the purchase category. Only 18 episodes (26.9%) contain a subsequent user turn within the same purchase-related mission, compared with 23 (34.3%) that contain any later user turn. Using conversation depth alone therefore overstates this follow-up availability by 27.8%. Across 47 retained user follow-up messages, no explicit post-advice purchase commitment, completed-purchase report, or purchase-category abandonment statement is observed. These zeros describe recorded statements, not conversion or persuasion rates. The paper supplies operational definitions, text-free annotations, and reproducible descriptive results. Its central finding is a measurement limitation: recommendation content is observable much more often than a buyer's subsequent decision. The selected historical sample, unvalidated AI annotations, and missing transaction outcomes do not support a population-level or causal estimate of persuasion.

Wed 9 SeptInformation Retrieval
The gist
This paper looks at conversations where AI assistants recommend products to people. The authors found that while AI assistants often suggest options, the conversations rarely record whether the person actually buys something or decides not to buy. In fact, follow-up messages from users after advice are uncommon, and no clear statements about buying or not buying appear in the records. The researchers highlight that it's hard to measure how much AI assistants influence buying decisions just from these chat logs.
Open 2609.09878v1

Speech models struggle with technical talk in science fields

$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants

Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplored, particularly in scientific areas. Scientific interactions introduce formidable challenges, involving rare technical terminology, spoken norms of abbreviations, and the natural verbalization of symbolic special expressions. In this paper, we introduce S$^3$-Bench, a systematic evaluation framework covering 10 major disciplines, consisting of a Knowledge set for speech question-answering and a Dialogue set for multi-turn progressive interactions with simulated user agents. By decomposing a complete atomic turn into stages of speech recognition, perception, knowledge utilization with reasoning, and response pronunciation, we systematically characterize the common challenges and performance tradeoffs of existing approaches. Furthermore, experiments on multi-turn interactions reveal persistent limitations in user adaptation and the generation of accurate, comprehensive, and efficient responses.

Wed 9 SeptComputation and Language
The gist
Voice assistants are good at general conversations but have trouble in scientific areas where people use complex terms and symbols. The authors created S3-Bench, a way to test how well these voice assistants handle science topics across 10 different fields. They look at speech recognition, understanding, reasoning, and speaking responses, finding that current models often fail to adapt well or give complete and accurate answers during longer talks. This shows that more work is needed to make voice assistants better for specialized scientific use.
Open 2609.09852v1

Large study reveals how people use multimodal AI task helpers

Large-Scale User Behavior Analysis in Multimodal AI-Assisted Manual Task Execution

Abstract: Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world tasks such as cooking and DIY through voice, text, image, and video interactions. Prior user studies have focused on controlled settings, leaving limited understanding of real-world CTA usage at scale. In this work, we present a large-scale study of CTA usage based on thousands of users in-the-wild. Our large-scale real-world data analysis unveils new understandings of (i) user-CTA interaction flows, (ii) user intents, (iii) user conversational traits, and (iv) behavioral factors associated with user satisfaction. Our findings reveal key opportunities for future research in CTAs, particularly in user interaction design and task engagement, concluding with concrete design guidelines.

Mon 7 SeptHuman-Computer InteractionArtificial IntelligenceComputation and Language
The gist
It can be hard to understand how people really use AI helpers for complex tasks like cooking or fixing things because studies often happen in limited, controlled places. The authors analyzed data from thousands of real users interacting naturally with a conversational AI that helps through voice, text, pictures, and video. They found patterns in how people talk to the AI, what they want to do, how they behave, and what leads to them feeling satisfied. These insights help improve the design of future AI helpers so they can be easier and more engaging to use.
Open 2609.07594v1