Grounded biographies improve question answering about long videos
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and LanguageInformation RetrievalMachine Learning
Summary
Answering questions about long videos is hard because the same object can appear many times and look similar to others. The authors created a method that links all appearances of a single object across a video into a ‘biography’ so the system can track it over time. This helps answer questions about events involving that object more accurately. Their approach worked better than previous methods on tests with day- and week-long videos.
What this means in practice
- •For video content platform developers: Enable video platforms to answer detailed questions about objects and events across hours of footage by linking object appearances into biographies.
- •For security monitoring teams: Improve tracking and investigation by associating individual entities through long surveillance videos for better incident understanding.
Authors
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua
Abstract
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.