Question Answering for Diagram-Rich Technical Meeting Videos
2026-07-11 • Software Engineering
Software Engineering
AI summaryⓘ
The authors developed LMVQA, a system that helps engineers find answers from long technical meeting videos by using AI to understand both speech and diagrams. It creates a searchable record of the video’s important parts, making it faster and cheaper to get accurate answers. Their tests showed big improvements in accuracy and speed compared to earlier methods, especially for videos with many diagrams. Engineers found LMVQA useful for finding specific information and understanding design decisions.
asynchronous communicationmultimodal question answeringlarge language modelstechnical meeting videosUML diagramsrequirements engineeringknowledge groundingtime-stamped evidence corpussoftware design rationalevideo indexing
Authors
Zhuoran Xu, Jia Li, Dayuan Tan, Mark Cole, Ish Ashraf, Sandeep Puri, Mehrdad Sabetzadeh, Shiva Nejati
Abstract
Software engineering increasingly relies on asynchronous communication artifacts, including recorded meetings where stakeholders discuss concerns, rationale, and decisions. These meetings often include diagram-based representations of requirements, system behavior, component interactions, and trace dependencies. Accessing knowledge from these meetings is challenging because recordings are long and relevant evidence is distributed across speech, slides, and technical diagrams. This paper reports our industrial experience developing and evaluating LMVQA, an LLM-based multimodal question-answering system for technical meeting videos. Developed in collaboration with engineers at Ciena, LMVQA supports the understanding of requirements and design intent by grounding answers in audio and visual evidence, with explicit handling of diagram-rich content such as requirements and UML diagrams. It processes each video once to build a reusable time-stamped evidence corpus for grounded question answering. Across a Ciena dataset and a public dataset, we show that LMVQA significantly improves answer accuracy compared to a state-of-the-art baseline, from 31% to 94% on the Ciena dataset and from 21% to 88% on the public dataset, with larger gains on diagram-rich videos. We further show that, after one-time indexing, LMVQA reduces average response time from 81.3s to 3.3s on Ciena and from 98.4s to 9.2s on the public dataset, while lowering average token-based LLM API cost by about 75%. Finally, our interviews with three domain experts show that engineers particularly value LMVQA for locating software-engineering-relevant information, revisiting rationale, and tracing answers to specific video segments.