Jazz combo audio dataset supports new music AI tasks

JazzSAMBA: A Synchronous and Asynchronous Multi-take Band Audio Dataset of Jazz Standards for Live Music Models

SoundArtificial IntelligenceMachine LearningMultimedia

Summary

Music AI systems have mostly focused on pop and rock, with little data available for jazz, which is special because musicians improvise a lot. The authors created JazzSAMBA, a collection of clean jazz recordings where each instrument is recorded separately, both live and overdubbed, and annotated with detailed musical information. This dataset covers many jazz standards played by professional musicians and includes audio, MIDI, and timing details. It helps improve music tasks such as separating instruments from a jazz band performance and generating accompaniment based on song charts.

What this means in practice

  • For music software developers: Build jazz-focused AI tools that separate instruments from live and overdubbed recordings using detailed dataset annotations.
  • For music tech startups: Create AI-powered apps that generate live jazz accompaniment conditioned on song charts and musician preferences.$Commercial implications: This dataset enables products that provide jazz accompaniment solutions with realistic instrument separation and chart adaptation.

Authors

Phillip Long, Jacob Nguyen, Jace Hosto, Gage Hosto, Jett Takazawa, Fares Nofal, Sebastian Stade, Nithya Shikarpur, Julian McAuley, Cheng-Zhi Anna Huang, Stephen Brade, Aleksandra Teng Ma

Abstract

Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-annotated corpus of clean per-stem combo recordings on standards. We introduce JazzSAMBA (Jazz Synchronous and Asynchronous Multi-take Band Audio) to fill this gap: the first originally recorded jazz-combo multitrack dataset of standards with asynchronous (overdubbed) and synchronous (live ensemble) protocols, preferred and alternate takes chosen by the musicians, and timed annotations for bars, chords, sections, and soloists. JazzSAMBA covers 76 standards by eight musicians on drums, bass, piano, trumpet, and saxophone, with per-stem audio, mixtures, and MIDI. It can support chart-conditioned accompaniment, combo source separation, and form-aware music information retrieval. We demonstrate the dataset on two tasks: a jazz combo source-separation baseline and a chart-conditioned accompaniment ablation. The dataset, code, and samples are linked from the project demo page.