Papers for

media accessibility teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio description generation optimized for timing and content choices

What, When, and How: Audio Description as Constrained Global Optimization

Abstract: Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.

Thu 24 SeptComputation and LanguageComputer Vision and Pattern Recognition
The gist
Making movies accessible to people who are blind means adding spoken descriptions of what’s happening on screen between dialogues. The authors show that this task isn’t just about describing what’s visible, but also about deciding what’s important to say, when to say it without overlapping dialogue, and how to say it briefly. They created a system that uses smart language models together with optimization techniques to pick and schedule descriptions for movie scenes. Their method improves on previous approaches in describing movies more thoughtfully and fitting descriptions better into the timing, although there’s still a gap compared to professional human describers.
Open → 2609.30121v1