AURA enables stepwise music track editing using multimodal dialogue integration
AURA: Unified Multimodal Framework for Conversational Music Editing
SoundArtificial Intelligence
Summary
Editing music piece-by-piece based on instructions is usually done for each request separately, which makes it hard to gradually improve a track over a conversation. The authors present AURA, a new system that understands the full history of conversation, plus optional pictures and example sounds, to figure out exactly what changes to make. It uses a combination of a language model that processes these inputs and a sound generation module that carefully edits music without disturbing parts that should stay the same. Tests show that AURA performs much better than current methods at making correct edits and keeping good parts intact.
What this means in practice
- •For music producers: Edit music tracks interactively by refining audio through ongoing conversations using text, images, and reference sounds for precise control without redoing full mixes.$Commercial implications: Enables creation of conversational music editing tools that improve workflow efficiency for professional producers and studios.
- •For podcast editors: Apply incremental audio edits to podcast tracks by interpreting dialogue-based instructions combined with example clips to keep undesired changes minimal.
Authors
Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha
Abstract
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaffected content. AURA optimizes only 91M parameters while retaining 1.9B frozen backbone parameters. Experiments on Slakh2100 and MoisesDB demonstrate substantial improvements in edit correctness and content preservation over existing instruction-guided methods, including a 4-5 times reduction in FAD for out-of-domain addition and removal.