Summary
When people work together, especially on tricky tasks, one person often helps another by giving instructions or physically assisting. Existing datasets mostly look at either a single person working or remote help, but do not fully capture the close interaction between a helper and a performer. This paper presents DYAD, a dataset that records synchronized video, audio, and task details from people working side-by-side on assembling a gearbox. The dataset links requests for help, the helper’s responses, and task progress, allowing researchers to study how help is sought and given in real time. The authors also show some example tasks to test parts of this interaction, like predicting when help will be needed or generating appropriate helper responses.
multimodal datasethuman assistanceegocentric sensingtask state trackinghelp seekingintervention choiceassembly taskHoloLenscausal step understandingmacro-F1 score
Authors
Akhil Ajikumar, Mahya Qorbani, Sakib Reza, Sean Andrist, Mohsen Moghaddam
Abstract
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They do not jointly link a co-located helper's verbal and physical interventions to performer requests, task state, assistance triggers, and outcomes. We introduce DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly. Across 20 sessions, one trained helper follows a guidance-first policy while assisting HoloLens 2 wearers. DYAD links 528 task-step intervals and 611 performer requests with 851 valid assistance records spanning verbal and physical help. DYAD's annotations span the assistance process; three reference tasks evaluate selected components rather than an end-to-end system: causal step understanding, pre-onset mode anticipation, and instructor response generation. On 829 eligible mode events, the strongest four-seed RGB mean is 0.548 +/- 0.007 macro-F1; causal metadata reaches 0.624 and a privileged trigger mapping 0.915, revealing information not recovered from pre-onset RGB. DYAD's contribution is not scale, but a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.