A Hybrid Ensemble Architecture for Symbolic and Acoustic Feature Fusion in AI-Based Music Generation

: Artificial Intelligence (AI) in music generation has seen significant progress, yet many existing models lack the ability to capture the full spectrum of musical features such as harmony, rhythm, and temporal dependencies. Addressing this gap, a hybrid ensemble framework is proposed by integrating DeepJ, REMI + RNN, and MuseGAN to enhance symbolic music generation. The system utilizes a publicly available dataset combining the Million Song Dataset, Spotify API features, and Last.fm user interactions, encompassing over 50,000 tracks and 1,500 audio samples across 15 genres. Through a pipeline involving data preprocessing, Mel-spectrogram conversion, and feature extraction (including MFCCs, Chroma, and spectral contrast), each model learns specialized musical components. The ensemble fuses outputs via heuristic alignment and a post-processing autoencoder for improved musical coherence. Experimental results demonstrate that the ensemble achieves superior performance with an R² score of 0.90, MAE of 0.15, and RMSE of 0.22 outperforming individual models like MusicRNN, DeepJ, and MuseGAN. These findings validate that combining symbolic, rhythmic, and generative representations leads to more expressive and accurate music synthesis. The proposed model holds promise for real-world applications in AI-assisted composition, music recommendation, and adaptive sound design.

Paper

Full text

PDF

A Hybrid Ensemble Architecture for Symbolic and Acoustic Feature Fusion in AI-Based Music Generation

Semantic Scholar · 2025

Abstract

: Artificial Intelligence (AI) in music generation has seen significant progress, yet many existing models lack the ability to capture the full spectrum of musical features such as harmony, rhythm, and temporal dependencies. Addressing this gap, a hybrid ensemble framework is proposed by integrating DeepJ, REMI + RNN, and MuseGAN to enhance symbolic music generation. The system utilizes a publicly available dataset combining the Million Song Dataset, Spotify API features, and Last.fm user interactions, encompassing over 50,000 tracks and 1,500 audio samples across 15 genres. Through a pipeline involving data preprocessing, Mel-spectrogram conversion, and feature extraction (including MFCCs, Chroma, and spectral contrast), each model learns specialized musical components. The ensemble fuses outputs via heuristic alignment and a post-processing autoencoder for improved musical coherence. Experimental results demonstrate that the ensemble achieves superior performance with an R² score of 0.90, MAE of 0.15, and RMSE of 0.22 outperforming individual models like MusicRNN, DeepJ, and MuseGAN. These findings validate that combining symbolic, rhythmic, and generative representations leads to more expressive and accurate music synthesis. The proposed model holds promise for real-world applications in AI-assisted composition, music recommendation, and adaptive sound design.

Similar papers

© 2026 NYSGPT2525 LLC