Papers for

podcast editors

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Music stem retrieval improved by separate slot embeddings for tracks

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

Abstract: Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.

Mon 28 SeptSound
The gist
When music producers want to find isolated parts of songs like a guitar or drums, they search large libraries of these parts called stems. The authors propose a new method named Stembed that breaks down a full song mix into several separate pieces called slot embeddings, each representing a possible stem. This method better matches parts of a mixture to individual isolated stems compared to previous approaches that looked at the entire mix as one single piece. It works even without knowing which instrument family to focus on and can predict how many stems there are.
Open → 2609.35672v1

AURA enables stepwise music track editing using multimodal dialogue integration

AURA: Unified Multimodal Framework for Conversational Music Editing

Abstract: Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaffected content. AURA optimizes only 91M parameters while retaining 1.9B frozen backbone parameters. Experiments on Slakh2100 and MoisesDB demonstrate substantial improvements in edit correctness and content preservation over existing instruction-guided methods, including a 4-5 times reduction in FAD for out-of-domain addition and removal.

Sun 13 SeptSoundArtificial Intelligence
The gist
Editing music piece-by-piece based on instructions is usually done for each request separately, which makes it hard to gradually improve a track over a conversation. The authors present AURA, a new system that understands the full history of conversation, plus optional pictures and example sounds, to figure out exactly what changes to make. It uses a combination of a language model that processes these inputs and a sound generation module that carefully edits music without disturbing parts that should stay the same. Tests show that AURA performs much better than current methods at making correct edits and keeping good parts intact.
Open → 2609.14344v1