Summary
Most music generation tools create a whole piece at once, but real music composition often happens through many small changes over time. The authors show that a type of large language model, originally designed to follow instructions in general, can be adapted to edit and improve music scores incrementally. They trained the model to understand specific editing actions like adding chords or changing keys, which lets it work on music step-by-step while keeping track of previous edits. Their results show this approach is technically possible, though it’s not yet clear if it makes better music or helps musicians work with AI.
What this means in practice
- •For music software developers: Build music editing tools that let users revise symbolic scores interactively with AI assistance guided by specific editing operations.
- •For game audio designers: Implement AI-based dynamic music systems that incrementally modify scores in real time to adapt to game events or player actions.
Authors
André Ricardo Ducca Fernandes, Jean-Pierre Briot, Simone Diniz Junqueira Barbosa1, Hélio Côrtes Vieira Lopes
Abstract
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.