Beyond the Mirror: Balancing Interaction Modality and Avatar Fidelity in Public 3D Virtual Try-On Systems

2026-08-24Human-Computer Interaction

Human-Computer Interaction
AI summary

The authors studied virtual try-on systems on large public screens that let people try clothes using gestures without touching anything. They found that the main cause of tiredness is the delay between moving and the system responding, not the gestures themselves, and improved this to make interactions feel easier and cleaner. They also discovered that very realistic avatars and mid-air gestures make users feel more present but can increase social awkwardness. Using simpler avatars helps people feel less embarrassed in public, and gestures help maintain trust in the experience even if the avatar looks less realistic. The authors suggest adjusting avatar realism based on context to balance privacy and immersion.

Virtual Try-On (VTON)3D AvatarMarkerless Motion CaptureVisuomotor LatencyMid-air InteractionVirtual EmbodimentVisual FidelitySocial InhibitionPsychological MaskHuman-Computer Interaction
Authors
Yueqian Guo, Tianzhao Li, Xin Lv
Abstract
Virtual Try-On (VTON) systems deployed on large public displays face a dual barrier: the physical strain of mid-air interaction and the social inhibition caused by public self-consciousness. This paper presents a real-time 3D avatar system integrating markerless motion capture with dynamic visual fidelity control to investigate and mitigate both barriers. Through a dual-study empirical evaluation, we first decoupled physical fatigue from gesture interaction ($N=20$), demonstrating that interaction fatigue is primarily driven by visuomotor latency rather than the physical act of gesturing; our optimized low-latency gesture pipeline achieved usability comparable to touchscreens while delivering superior immersion and hygiene. Building on these insights, our second study ($N=25$) investigated the "avatar fidelity paradox" via a $2 \times 2$ factorial design manipulating interaction modality (gestures vs. touch) and visual fidelity (photorealistic MetaHuman vs. stylized mannequin). Results reveal that while high fidelity and mid-air gestures independently maximize virtual embodiment ($p < .05$), their combination elicits the highest social awkwardness. Crucially, low-fidelity avatars serve as a "psychological mask" that alleviates public embarrassment during expressive gestures, while mid-air gestures simultaneously act as a compensatory mechanism to preserve perceived try-on trust despite reduced visual realism. Finally, we propose a context-aware fidelity framework to balance privacy, immersion, and commercial trust in public spatial interactions.