A Distributed Control Approach for Multi-Agent Reinforcement Learning in Collaborative AI Music Performance

Automatic harmony generation is a critical challenge in AI-driven music composition, requiring simultaneous preservation of tonal coherence in harmonic structures, controllable emotional expression, and artistic integrity of the final piece. Addressing limitations in existing methods—such as inadequate emotional control, lack of interpretable musical semantics, and absence of harmonic norms—this paper proposes a hierarchical harmony generation framework based on Inverse Reinforcement Learning (IRL) and Deep Reinforcement Learning (DRL). First, IRL is employed to learn latent harmonic feature representations from expert harmonic data, constructing an interpretable harmonic semantic space. Second, DRL dynamically adjusts harmonic color and tension development by using harmonic affect curves as policy reward signals. Finally, a harmonic refinement module grounded in compositional theory ensures tonal establishment and the integrity of cadential structures. Experimental results demonstrate that this method achieves emotional consistency and structural coherence while preserving artistic expressive flexibility. It also receives high recognition in subjective listening evaluations and editability analyses.

Paper

Full text

PDF

A Distributed Control Approach for Multi-Agent Reinforcement Learning in Collaborative AI Music Performance

Semantic Scholar · 2025

Abstract

Automatic harmony generation is a critical challenge in AI-driven music composition, requiring simultaneous preservation of tonal coherence in harmonic structures, controllable emotional expression, and artistic integrity of the final piece. Addressing limitations in existing methods—such as inadequate emotional control, lack of interpretable musical semantics, and absence of harmonic norms—this paper proposes a hierarchical harmony generation framework based on Inverse Reinforcement Learning (IRL) and Deep Reinforcement Learning (DRL). First, IRL is employed to learn latent harmonic feature representations from expert harmonic data, constructing an interpretable harmonic semantic space. Second, DRL dynamically adjusts harmonic color and tension development by using harmonic affect curves as policy reward signals. Finally, a harmonic refinement module grounded in compositional theory ensures tonal establishment and the integrity of cadential structures. Experimental results demonstrate that this method achieves emotional consistency and structural coherence while preserving artistic expressive flexibility. It also receives high recognition in subjective listening evaluations and editability analyses.

Similar papers

© 2026 NYSGPT2525 LLC