Papers for

ai assistant developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Stashbird cuts memory tokens for ai agents with indexed updates

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Abstract: AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

Mon 28 SeptArtificial IntelligenceMachine Learning
The gist
AI assistants need to remember past conversations accurately, including who said what and when information changes. The authors created Stashbird, a system that organizes memory with clear links to original conversations and allows easy updates or deletions. This system uses much fewer resources to process and recall information than some existing methods, while maintaining similar or better accuracy. Stashbird is tested across multiple benchmarks involving long-term conversational memory.
Open → 2609.34242v1

Agent behaviors on tasks reveal different styles beyond success

AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks

Abstract: Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users' preferences and needs. For example, agents differ in whether they ask clarifying questions or search the web. We introduce HABIT, a taxonomy of 23 behavioral axes in five categories, which three authors and three LLMs derive bottom-up from 408 agent trajectories across 17 domains. On held-out tasks, HABIT distinguishes models more clearly than existing taxonomies of human values and agent actions while supporting comparably consistent annotation. Building on HABIT, we construct AgentHABIT, a benchmark that profiles each agent's behavioral tendencies from its trajectories on 86 everyday tasks. Profiling 18 models with AgentHABIT reveals a range of distinctive tendencies. For example, most GPT and Claude models state their assumptions and offer alternatives when requirements conflict, whereas Qwen and Google's models more often leave assumptions or changes to requirements unstated. These profiles remain recognizable even when built from entirely different sets of tasks, indicating that they reflect general tendencies rather than task-specific behavior. Prompting agents to adopt specific behaviors shifts some axes readily but barely changes others, while fine-tuning on another model's trajectories changes only part of a model's profile and leaves much of it intact. Overall, HABIT and AgentHABIT provide a systematic framework for characterizing how agents carry out everyday tasks beyond task success, offering insights to guide the development of agents whose behavior better fits users' needs.

Sat 26 SeptArtificial Intelligence
The gist
When AI assistants complete everyday tasks, they often do so in many different ways that may not always match what people prefer. The authors created a detailed checklist called HABIT to describe 23 types of behaviors that these AI agents show, like asking questions or searching online. They tested 18 agents on many tasks and found clear patterns in how each tends to behave. This helps us understand AI assistants better and could guide making them act more like people want.
Open → 2609.32795v1

Data artists develop stories and visuals together through sketches and prototypes

Understanding Creative Design Practices among Data Artists

Abstract: Creative and artistic data visualizations communicate stories and invite engagement, yet how designers develop their expressive forms remains poorly understood. We investigate the design process of data artists through three complementary studies: an analysis of 40 public project accounts, artifact-anchored interviews with seven experienced data artists, and a design task-based study with eight data artists. We find that stories and visual forms develop together through data exploration, reference adaptation, sketching, and prototyping. Inspiration comes from data, existing work, everyday imagery, and personal experience. Sketches help develop mappings, while prototypes with real data can reshape representations and intended stories. Designers consider multiple possibilities but typically develop one direction at a time, partly because producing alternatives is costly. Their choices balance meaning, visual appeal, readability, and feasibility. These findings inform tools connecting stories, references, and real data, including AI assistance that supports testing and revising alternatives while preserving designers' creative judgment.

Wed 23 SeptHuman-Computer Interaction
The gist
Creative data visualizations tell stories and invite people to explore information, but how artists design these visuals is not well known. The authors studied data artists through projects, interviews, and design tasks, finding that artists combine exploring data with sketching and prototyping to shape their stories and images. They get ideas from data, other artworks, everyday sights, and personal memories. Artists usually focus on one design direction at a time since making multiple options is costly, balancing meaning, beauty, clarity, and what they can do.
Open → 2609.28715v1

FlowSeal stops privacy leaks in AI assistants using outside control

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Abstract: Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular party. Existing defenses address this by making the agent's backend LLM more privacy-preserving through stronger system prompts, training, or explicit consent-checking procedures, but this approach has a structural challenge: whenever enforcement is a judgment the LLM makes over the same conversational context an adversary controls, the enforcement mechanism and the attack surface coincide. We demonstrate this against existing defenses with three new attacks that require only ordinary agent interaction and no prompt injection: Collaborative Workspace Lure reframes an extraction attempt as collaborative work; Semantic Obfuscation Attack induces disclosure through omission rather than through anything the agent writes; and Channel Decoupling Attack splits the extraction request and the disclosure across independent channels. All three achieve substantially higher leak rates than the attacks these defenses were originally designed to withstand. Guided by this observation, we present FLOWSEAL, a defense that enforces confidentiality through a tool-level interceptor outside the LLM's context, grounded in data provenance and an information-flow-control lattice with controlled declassification. Evaluated across three benchmarks, five prompt-based baselines, and eight attacks, including a real agent executing live tool calls through MCP, FLOWSEAL reduces leak rates to near zero (e.g., 52.2% to 0.5% against Collaborative Workspace Lure) while preserving task utility, regardless of the underlying LLM backend.

Sat 12 SeptCryptography and SecurityArtificial Intelligence
The gist
Personal AI assistants that use large language models (LLMs) often need to access sensitive user information, creating risks that private data might be leaked. The paper's authors show that existing privacy protections can be bypassed by clever tricks that confuse the AI inside the same conversation. They introduce FlowSeal, a new defense that controls data flow outside the AI’s internal context, using careful tracking of information and limits on what can be shared. Tests show FlowSeal greatly reduces leaks while still allowing useful tasks to be done, no matter which LLM powers the assistant.
Open → 2609.14003v1