Generative tutorial improves physical task guidance with live visuals

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Human-Computer InteractionArtificial Intelligence

Summary

Visual instructions for handmade tasks often don't match the user's workspace, making them hard to follow. The paper introduces a system that creates live images and videos showing exactly what to do in your own space. This approach helps people do tasks better and faster, and makes instructions feel more trustworthy. The system was tested with 24 people who performed better compared to traditional pre-made guides.

What this means in practice

  • For home improvement contractors: Use live-generated visual instructions that fit their specific work environment to improve accuracy and confidence during physical installations.
  • For assembly line supervisors: Deploy augmented reality guidance that dynamically adapts tutorials to the workspace for better worker performance and reduced errors.$Commercial implications: Enables sale of context-aware AR instruction systems that improve factory assembly workers’ task quality and speed.

Authors

Muzhe Wu, Zuchen Li, Xu Wang, Anhong Guo

Abstract

Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.