Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments

2026-07-20Sound

Sound
AI summary

The authors study how pianists can tailor recordings of orchestra accompaniments to fit their solo playing without needing written music scores. They use three audio types: solo piano, orchestra-only, and mixed recordings, with the mixed ones helping to line up the solo and orchestra parts. To better synchronize these parts despite differences in sound, the authors develop Dense-Sparse DTW, a method that focuses on key timing cues for alignment. Testing on several piano concertos, their method performs as well or better than other complex sound separation techniques.

Music Minus OneDynamic Time Warpingaudio alignmenttime-scale modificationspectral mismatchsource separationpiano concertoaudio synchronization
Authors
TJ Tsai, Kavi Dey, Yigitcan Ozer, Meinard Muller
Abstract
In this study, we explore how pianists can customize Music Minus One (MMO) concerto accompaniments to match their playing style. Bypassing the need for a symbolic score, often not available digitally, we use three types of audio data: solo piano recordings, MMO orchestra-only recordings, and mixed recordings of both piano and orchestra (e.g., from YouTube). The mixed recording serves as an intermediary reference to align the solo and orchestra parts, with only the orchestral part being adjusted through time-scale modification to synchronize with the user's playing. The main challenge with estimating these alignments is the spectral mismatch between recordings containing different musical parts. Motivated by this application scenario, we introduce Dense-Sparse DTW, a variant of Dynamic Time Warping (DTW) that is designed to improve robustness of alignments to spectral mismatch by focusing on aligning a selected subset of audio frames containing prominent timing cues. We collect and annotate data from four piano concerto movements and establish a framework for generating and evaluating customized accompaniment recordings. On this benchmark, we show that Dense-Sparse DTW has better or comparable performance than more complex approaches based on source separation and spectral subtraction techniques.