Papers for

ai product designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Users value conditions set around AI agent use more than outcomes

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Abstract: Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect's corpus share, values clustered not at the agent's outputs but at the operating conditions users set around a run. Values were usually met where users described what the agent delivered, in five of six groups, and mostly unmet where users described supervising it, in all six groups. We conceptualize this pattern as value-sensitive delegation. Supporting human values requires attention not only to what an agent accomplishes, but to the conditions users set around delegation, including cost, access, and oversight.

Fri 18 SeptHuman-Computer InteractionArtificial Intelligence
The gist
People often let AI agents do tasks for them, but usually we only check if the task was finished. The authors studied lots of user posts about an AI agent called OpenClaw to see what values users cared about. They found that users cared more about how they set up and monitored the AI, like costs and fairness, rather than just what the AI delivered. This means creating AI tools also means paying attention to how people control and oversee them.
Open 2609.22067v1

Ai companions show racial traits that affect how people see them

Stereotypically Yours: Portrayal and Perception of Race-Coded AI Companions

Abstract: AI companions can purportedly adopt racial personas, raising questions about how they represent identity and how users interpret these portrayals. We combined an algorithmic audit of race-coded AI personas with interviews with 12 companion users who interacted with a probe. Our audit revealed systematic differences, such as Asian-coded male personas receiving higher submissiveness scores than White counterparts, and Black, Hispanic, and Indigenous male personas receiving higher aggression scores than their White counterparts in open-weight models. Interviews revealed that participants envisioned AI companions as offering cultural familiarity and outside perspectives, but differed in which portrayals they considered meaningful or stereotypical. Some rejected overt racial signaling while still expecting culturally distinctive responses. Triangulating these findings with theory, we highlight how social norms and cultural expectations complicate efforts to support meaningful racial representation without reproducing stereotypes. We discuss how companion personalization should be evaluated beyond user satisfaction to account for broader representational harms.

Thu 17 SeptHuman-Computer Interaction
The gist
AI companions can take on traits linked to different races, and this changes how users feel about and understand them. The authors found that AI personas coded as Asian men acted more submissive, while those coded as Black, Hispanic, or Indigenous men showed more aggression compared to White personas. People who used these AI companions had different opinions about how much race should show up—some wanted cultural connections without obvious stereotypes. The authors say that making AI companions truly respectful of race is hard because social ideas and stereotypes get in the way.
Open 2609.20637v1

Personalised language models influence user trust and sharing habits

Tailored to you: longitudinal effects of personalising language models

Abstract: Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.

Thu 17 SeptArtificial Intelligence
The gist
People use personalised language models in different ways depending on how the models remember or learn about them. The study found that repeated use of any language model changes how people interact with it, but personalisation can lead to more or less sharing of personal details and feelings about privacy. For instance, models that remember past conversations make users share more and feel less creeped out. On the other hand, models personalised from upfront surveys can make users regret sharing personal info. The authors highlight that different ways of personalising affect users in complicated ways.
Open 2609.20077v1

People increasingly trust large language models for personal advice

LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions

Abstract: We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of AI reliance at scale and understand how people are offloading judgment and decision-making to AI. Applying our typology to public usage data (68K prompts from WildChat and ThoughtTrace), we find that LLM-as-oracle use has increased over time (2023-2026) and is more prevalent among younger users. We further build a privacy-preserving data donation tool to analyze individuals' longitudinal usage data (140K prompts from 52 participants), identifying similar trends. People are often unaware of their own LLM-as-oracle use, and express dissatisfaction with this behavior after seeing our tool's analysis. Finally, we identify two drivers of LLM-as-oracle use: people's perceptions of AI and the behavior of AI models themselves, which motivate possible interventions to support users' self-deliberation.

Sun 13 SeptComputers and SocietyArtificial IntelligenceComputation and Language
The gist
People are starting to rely on AI language models like ChatGPT for answers to personal, subjective questions as if the AI knows everything. The authors studied how and why people use these models as all-knowing oracles, finding that younger people do this more often and that people often don’t realize they are doing it. They created tools to analyze AI usage while protecting privacy and found users sometimes feel unhappy about depending on AI for personal decisions. The authors also identified factors that make people rely on AI like their beliefs about AI and how the AI responds.
Open 2609.14849v1

Teens rely heavily on ai chatbots but face emotional challenges

A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots

Abstract: AI companions provide socially engaging interaction through availability, personalization, memory, roleplay, and emotionally responsive language. For teens, these systems may support sensitive self-disclosure, identity exploration, and relationship rehearsal while shaping intimacy expectations, offline relationships, emotional wellbeing, and self-understanding. We analyzed 17,053 verified quotations from 3,930 teen-relevant Reddit posts using thematic analysis. We identified 53 topics across seven thematic groups. Users described AI companions as sources of comfort, recognition, identity exploration, and relationship rehearsal, but also reported problematic attachment, social substitution, emotional dependence, and disruption to academic and social life. Roleplay, memory, perceived reciprocity, unwanted romantic or sexual role drift, privacy concerns, platform changes, and service interruptions shaped users' boundaries and control. Awareness that the AI was artificial did not prevent guilt, obligation, grief, or distress. These findings show that companion-AI safety must address relationships over time through user-controlled memory, privacy, relational boundaries, and healthy disengagement.

Sun 13 SeptHuman-Computer InteractionArtificial Intelligence
The gist
Teens use AI chatbots as friends to share feelings, explore who they are, and practice relationships. The authors studied thousands of online teen posts to see how these chatbots affect moods and friendships. While these chatbots offer comfort and identity support, some teens become too attached or replace real social interactions. Problems like emotional dependence, unwanted romantic conversations, and privacy worries also came up. The authors suggest safer chatbot designs that give users control over memory, privacy, and how they end relationships.
Open 2609.14843v1

Challenges arise in applying existing design rules to AI companions

Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study

Abstract: With the rapid proliferation of large language model (LLM)-based systems, AI companions have emerged as conversational agents designed to cultivate emotional connection rather than primarily to support humans in instrumental tasks. Because engagement with AI companions involves relational, emotional, and potentially long-term interactions, their design is consequential. Prior work has offered guidance for designing trustworthy and relational AI systems and has begun to examine design for AI companionship. However, while such work provides insights into possible design solutions, less is known about what makes AI companion design difficult as a design problem. To examine this challenge, we assessed the applicability of existing design recommendations from adjacent domains in the context of AI companion design. Our multi-method investigation unfolded across four phases: literature review, practitioner co-analysis, internal heuristic evaluation, and external expert assessment. Throughout this process, we synthesized nine design principle areas that surfaced tensions in the applicability of existing recommendations to AI companion design. Our findings show that ethical and UX-oriented considerations are deeply intertwined and often require context-sensitive application. We document a systematic, multi-method problem analysis that uses these principle areas as an analytic artifact to examine why existing recommendations cannot be directly transferred to AI companion contexts.

Sun 13 SeptHuman-Computer InteractionArtificial Intelligence
The gist
AI companions are conversational programs meant to create emotional bonds with people over time. Designing these AI companions is tricky because existing design advice from related areas doesn’t always fit well. The authors studied how current guidelines apply to AI companions by reviewing literature, working with designers, and evaluating expert feedback. They found that ethical and user experience issues are closely linked and depend heavily on context. This means new or adapted design principles are needed specifically for AI companions.
Open 2609.14236v1

Human and AI qualities define great co-workers in AI workplaces

What Makes a Great Co-Worker in an AI-Native Workplace?

Abstract: As knowledge work grows interdependent between humans and AI, we ask what makes a great co-worker in an AI-native workplace. To answer this, we conducted 22 interviews and a large-scale mixed-methods survey of 1,534 knowledge workers at a multinational technology company. We contribute BACI, a framework of 75 co-worker qualities that apply to humans and AI, spanning Benevolence, Ability, Cooperativeness, and Integrity. Comparing priorities for humans and AI identified 11 co-worker archetypes and revealed disagreement over whether AI should have warmth, take initiative, or own outcomes. We also show how priorities for these archetypes varied with workers' individual characteristics. Lastly, we contribute a taxonomy of AI work etiquette capturing the obligations co-workers expect of one another when preparing, sharing, and taking responsibility for AI-supported work. Based on these findings, we derive implications to inform worker-centric AI and workplace design.

Sat 12 SeptHuman-Computer Interaction
The gist
As people work more closely with AI in their jobs, this study explores what makes a great co-worker when both humans and AI are involved. The authors interviewed workers and surveyed over 1,500 employees at a large tech company to identify important qualities shared by humans and AI, such as kindness and ability. They found different opinions about how AI should behave, like whether it should show warmth or take initiative. They also created guidelines for how humans and AI should work together respectfully and responsibly.
Open 2609.13786v1

Reasoning formats affect how people check AI model answers

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

Abstract: Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.

Tue 8 SeptMachine LearningHuman-Computer Interaction
The gist
People need ways to understand and trust answers from large language models (LLMs). The authors studied six different ways to show the reasoning behind LLM answers to see which best helps people check and trust those answers. They found that while many people liked clear, structured plans, the simplest step-by-step reasoning helped them spot errors and decide whether to trust the answer better. Some preferred formats made people too confident or triggered unnecessary doubts, showing a mismatch between what users like and what works best for verifying AI outputs.
Open 2609.09038v1

Sustained AI companion use links to lower well-being through less human contact

Living with AI Companions: Sustained AI Companionship Predicts Lower Well-Being Through Lower Human Interaction

Abstract: AI chatbots are increasingly used for companionship, emotional support, and personal self-disclosure; however, how social engagement with these systems unfolds over time and shapes users' well-being remains unclear. To address this, we conducted a two-wave longitudinal study of CharacterAI users, surveying 1,182 participants at baseline and 439 after a mean follow-up of 12 months. We examined how social engagement with AI companions evolves and how these longitudinal engagement patterns may influence well-being through two hypothesized pathways: sustained social engagement over time and the displacement of human social interaction. We found that interaction intensity, companionship use, and self-disclosure all showed substantial continuity over time. Greater interaction intensity at baseline predicted greater subsequent interaction intensity, companionship use, and self-disclosure. Consistent with the longitudinal engagement pathway, sustained social engagement across these dimensions was consistently associated with lower well-being. Results further support the social displacement pathway, indicating that these links were mainly explained by lower in-person social interaction. These findings highlight the importance of designing AI companions that support human social relationships without displacing them

Mon 7 SeptHuman-Computer Interaction
The gist
People use AI chatbots for emotional support and companionship, but this study found that those who keep using these AI companions a lot over time tend to feel less well. The researchers show that this effect seems to happen because interacting more with AI companions often means spending less time with real people. The study suggests designers should think about ways AI companions can support human relationships rather than replace them.
Open 2609.07243v1