AI summaryⓘ
The authors explored whether a large language model (LLM) can learn what kind of symbolic music a person likes just from rankings of music pieces. They had the LLM guess what the listener values, create new music, get rankings, and then explain its guesses to improve over time. While this method did not generally beat just creating varied music without feedback, it did help for some harder-to-reach music preferences. The study also found that understanding and describing these preferences depends heavily on the LLM's ability to recognize and talk about musical features. Overall, the authors show that using only rankings to guide music generation is possible but has clear limitations tied to the model’s capabilities.
generative music systemlarge language modelsymbolic music generationin-context learningABC notationpreference rankingvalue criteriatransfer learningfeedbackmixed-effects modeling
Authors
Futa Hidaka, Naomi Imasato, Kazuki Miyazawa, Takato Horii
Abstract
Adapting a generative music system to an individual's taste requires learning what that listener values. Listeners can rank pieces, but their underlying criteria may be tacit and difficult to articulate. We ask whether and under what conditions a large language model (LLM) can adapt symbolic music generation from rankings alone and construct transferable natural-language descriptions of value criteria. In our iterative in-context learning framework, the LLM formulates hypotheses, generates candidate pieces in ABC notation, receives a ranking, and periodically infers and verbalizes value criteria from history to guide later generation. We evaluate the framework against 16 simulated raters in 480 adaptation runs using mixed-effects modeling, an ablation, and transfer tests on unseen music. Overall, the framework did not outperform a feedback-free diverse-generation baseline, but did so for two value functions with targets difficult to reach through simple sampling. How atypical the target was relative to the LLM's feedback-free generation tendencies predicted adaptation difficulty. Moreover, higher value during adaptation did not imply identification of the criterion as a general rule. On unseen music, acquired descriptions and histories improved generation for more value functions than they improved preference prediction, which remained near chance. Some gains were associated with acoustic proximity to music in the context, but others were not. These findings show that rankings alone can guide generation under limited conditions, while transferable criterion inference remains constrained by the foundation model's ability to recognize, reason about, and verbalize musical attributes.