Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents

2026-07-20Human-Computer Interaction

Human-Computer Interaction
AI summary

The authors studied how computer agents perform tasks using GUIs and noticed that current feedback is mostly text and hard to follow. They created a system called Sidekick that uses sounds, visuals, and speech to show what the agent is doing in different ways depending on whether it is running quietly, just finished, or working in front of you. In a test with 30 people, Sidekick helped users multitask better and kept them more aware of the agent's progress and actions than old text-based methods. The authors show examples of Sidekick working well and discuss how it might improve long-term teamwork between humans and agents.

Computer Use AgentsGraphical User InterfaceMultimodal FeedbackMultitaskingAgent TransparencyContext ResumptionAmbient CuesHuman-Agent CollaborationUser Interface DesignError Traceability
Authors
Ruei-Che Chang, Wenqian Xu, Dingzeyu Li, Bryan Wang, Anhong Guo
Abstract
Computer Use Agents (CUAs) can autonomously execute complex, multi-step tasks within GUIs, enhancing efficiency through parallel multitasking. However, our formative studies with CUA experts and GenAI users indicated that current feedback is primarily text-based, requiring sustained attention to monitor progress and offering limited visibility to trace past GUI interactions. Based on the findings, we developed a prototype system, Sidekick, for communicating CUAs' status with multimodal feedback across different stages of interaction: (i) When CUAs run in the background, Sidekick signals its execution state through ambient cues. (ii) Upon resuming interaction with CUAs, Sidekick provides multimodal summaries of completed actions to support rapid context resumption. (iii) When CUAs operate in the foreground, Sidekick enhances transparency by verbalizing and visualizing the agent's reasoning. A study with 30 participants demonstrated that Sidekick significantly improved multitasking performance with CUAs compared to baseline systems that presented textual feedback either in a typical chat or in an ambient display. Sidekick supported progress awareness, and error and action traceability more effectively. Finally, we demonstrate the promise of Sidekick through several example applications, and discuss implications for long-horizon human-agent collaboration.