You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction

2026-08-03Human-Computer Interaction

Human-Computer Interaction
AI summary

The authors explain that augmented reality (AR) interfaces powered by AI change what info they show based on real-time situations, unlike traditional fixed interfaces. Because these AR interfaces adapt over time, the authors say we should evaluate them by looking at how they perform across many different situations and over long periods, not just in one test. They also note current studies mostly test these interfaces in simple, one-time settings and don’t capture how users handle multiple tasks or how their trust changes over time. To improve this, the authors suggest new ways to measure performance that consider task interference, changing contexts, and long-term user trust.

augmented realityAI-mediated interfaceheads-up displayoptical see-through head-mounted displaycognitive loadperformance metricstask interferencelongitudinal studyuser trustcontext adaptation
Authors
Nuwan Janaka, Runze Cai, Yang Chen, Chenyu Zhao, Shengdong Zhao
Abstract
AI-mediated heads-up augmented reality (AR) replaces fixed interfaces with dynamically adapting ones that decide what information to present, in what form, and when, based on a continually changing context that cannot be fully anticipated beforehand. Although it remains an interface, its behavior over time is only partially specified at design time. We argue that this shift requires a corresponding change in evaluation: from snapshots to trajectories. A fixed interface is evaluated in a snapshot --- one context, one session, one set of task-performance metrics. A fluid interface must be evaluated over a trajectory --- a sequence of contexts with transitions, sampled from the distribution the interface will actually encounter, and tracked long enough for user trust to form, evolve, and potentially deteriorate. Drawing on the literature for heads-up AR multitasking enabled by optical see-through head-mounted displays (OST-HMDs), we find that current evaluation practice remains largely snapshot-based. Most studies use fixed-condition, single-session designs; interference between concurrent tasks is rarely quantified directly; and commonly used workload measures cannot disentangle cognitive load attributable to individual tasks. To address these limitations, we argue for three shifts: from isolated metrics to Performance Operating Characteristic (POC) interference frontiers, from fixed conditions to evaluation over context trajectories, and from single-session snapshots to longitudinal trust measurement.