Vocal Music under Phoneme-Conditional Analysis

2026-08-31Sound

SoundComputation and Language
AI summary

The authors studied how different languages sound when people sing without instruments. They wanted to see if specific speech sounds (phonemes) make each language's singing unique. By comparing similar parts of songs that include these special sounds to parts without them, they measured differences in the music’s acoustic features. Their method could correctly guess the language of a song 85.5% of the time by looking at these patterns. This suggests that the sounds of a language influence how its songs are sung.

phonemeacoustic analysisvocal musictypologysyllablelanguage identificationbalanced accuracyphonological structuremelodygenre
Authors
Hayoon Kim, Kyogu Lee
Abstract
The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are measurable and traceable to specific phonemes. To tackle this question, we introduce phoneme-conditional analysis, which isolates the acoustic effect of typologically distinctive phonemes by comparing marker syllables against matched non-marker controls within the same song, holding singer, melody, and genre constant. Across nine typologically diverse languages and thousands of songs, we measure effects along five acoustic dimensions. Song-level profiles built from these effects identify the language of an unaccompanied vocal at 85.5% balanced accuracy in a nine-way classification with folds grouped by artist; whether the separability arises by accumulation of the phoneme-local effects themselves is left open. Our findings suggest that phonological structure leaves systematic and measurable traces in how each language is sung.