Papers for
music streaming services
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Greek automatic lyric transcription improves with adapted Whisper models
Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation
Abstract: Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, where no prior benchmark for ALT exists. We present the first controlled study of Whisper adaptation for Greek ALT, investigating model scaling effects, task composition via multitask training in transcribe-translate ratios, and two-stage speech-to-singing adaptation. We also curate a segment-level aligned singing dataset based on the Greek Audio Dataset (GAD) using source separation and CTC forced alignment. Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models. The 2-stage adaptation in Whisper Large-v3 achieves a Word Error Rate (WER) of 27.2%, a significant improvement over zero-shot baselines, establishing the first Greek ALT benchmark.
Matching songs by listening habits beyond artist and genre
Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data
Abstract: This report presents results from Project Qualia, an ongoing effort to determine whether experiential similarity between songs, a structure not captured by genre or metadata taxonomies, can be recovered from real listening behavior. We constructed a large-scale dataset of listening sessions, comprising 1.29 billion scrobbles collected from 9,396 users via the Last.fm API and reduced through a preprocessing pipeline to 531.6 million training scrobbles across 28.6 million sessions. On this corpus, we trained a skip-gram Word2Vec model (Song2Vec), treating each session as a sentence and each track as a token. As anticipated, the resulting embedding space was dominated by artist identity, a consequence of single-artist runs within sessions. To test for a subtler, artist-independent signal, we developed an artist-residual procedure: subtracting each artist's centroid from its tracks' embeddings and evaluating whether the remainder retained structure. Mean cross-artist cosine similarity fell from 0.2487 in raw embedding space to 0.0005 in residual space, yet 4,577 cross-artist track pairs retained cosine similarity $\ge 0.70$ in residual space, forming coherent genre- and era-based clusters, including trip-hop, 1990s grunge, 2020 mainstream pop, and cross-composer classical piano pairs at cosine similarity up to 0.95. These results confirm that the training data contains experiential structure independent of artist identity, establishing an empirical basis for an architecture designed to learn this experiential layer directly.
Irish traditional tunes show diversity grows despite selection pressures
Population Ecology of Tunes
Abstract: How cultural repertoires maintain diversity under selection is a fundamental question in cultural evolution. We address this using thirteen years of weekly popularity data for approximately 20,000 Irish traditional tunes, fitting ecological birth-process models under neutral, frequency-dependent, and per-tune selection hypotheses. We find strong evidence that tunes differ in intrinsic fitness - some are systematically more likely to be learned than others. We find that 29% of the variance in fitness can be explained by a mixture of social and melodic features. Some tunes appear to be carried along via linkage due to the tradition of playing tunes in sets, analogous to selective sweeps in genetics. By measuring changes in fitness over time and comparing this with recordings we precisely identify the mechanism by which a long-dormant tune can become fit through a popular recording. Despite the directional selection, repertoire diversity increases, driven by the continual arrival of new compositions. These results demonstrate that selection and diversity can coexist in a cultural ecosystem, and establish Irish traditional music as a quantitatively tractable system for studying the evolution of cultural variants and understanding what makes a tune stand out.