A Distributed Control Approach for Multi-Agent Reinforcement Learning in Collaborative AI Music Performance
Automatic harmony generation is a critical challenge in AI-driven music composition, requiring simultaneous preservation of tonal coherence in harmonic structures, controllable emotional expression, and artistic integrity of the final piece. Addressing limitations in existing methods—such as inadequate emotional control, lack of interpretable musical semantics, and absence of harmonic norms—this paper proposes a hierarchical harmony generation framework based on Inverse Reinforcement Learning (IRL) and Deep Reinforcement Learning (DRL). First, IRL is employed to learn latent harmonic feature representations from expert harmonic data, constructing an interpretable harmonic semantic space. Second, DRL dynamically adjusts harmonic color and tension development by using harmonic affect curves as policy reward signals. Finally, a harmonic refinement module grounded in compositional theory ensures tonal establishment and the integrity of cadential structures. Experimental results demonstrate that this method achieves emotional consistency and structural coherence while preserving artistic expressive flexibility. It also receives high recognition in subjective listening evaluations and editability analyses.
Paper
Full text
A Distributed Control Approach for Multi-Agent Reinforcement Learning in Collaborative AI Music Performance
Semantic Scholar · 2025
Abstract
Automatic harmony generation is a critical challenge in AI-driven music composition, requiring simultaneous preservation of tonal coherence in harmonic structures, controllable emotional expression, and artistic integrity of the final piece. Addressing limitations in existing methods—such as inadequate emotional control, lack of interpretable musical semantics, and absence of harmonic norms—this paper proposes a hierarchical harmony generation framework based on Inverse Reinforcement Learning (IRL) and Deep Reinforcement Learning (DRL). First, IRL is employed to learn latent harmonic feature representations from expert harmonic data, constructing an interpretable harmonic semantic space. Second, DRL dynamically adjusts harmonic color and tension development by using harmonic affect curves as policy reward signals. Finally, a harmonic refinement module grounded in compositional theory ensures tonal establishment and the integrity of cadential structures. Experimental results demonstrate that this method achieves emotional consistency and structural coherence while preserving artistic expressive flexibility. It also receives high recognition in subjective listening evaluations and editability analyses.