From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
2026-08-25 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors review the development of smart glasses as devices that combine what a person sees, hears, and does to provide useful real-time assistance. They point out that current research is fragmented and that the real challenge is building systems that work reliably over time and can be controlled and corrected. To address this, the authors propose a unified framework to study smart glasses, organizing their features, tasks, and applications into structured categories. They also introduce ways to evaluate and compare smart glasses in realistic settings, aiming to make them more trustworthy and practical.
smart glassesfirst-person visionaugmented realityhuman-computer interactionegocentric perceptioncontextual assistanceembodied intelligenceevaluation protocolslongitudinal validation
Authors
Jiangning Zhang, Haojun Chen, Yong Liu
Abstract
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.