Synthesis of Disparate Audio Species via Recurrent Neural Embedding

Sound synthesis plays a crucial role in the alteration of style, mood, and auditory affordances embedded in sounds. However, currently, sound synthesis has been limited mostly to a single species of sound, music sound. And, the synthesis of multiple disparate sound species that can transform the sound synthesis into another dimension is yet to be established. In this study, we present a method for synthesizing disparate sound species, in particular, music and natural sounds in real-time. We achieve this by interpolating the timbre embedding of music and natural sounds together with independent pitch conditioning, where a bidirectional LSTM stack is used to encode a full sequence of music and natural sounds while disentangling pitch from timbre. The effectiveness of the proposed sound synthesis is validated by evaluating the R2 score and the Structural Similarity Index Measure (SSIM) associated with synthetic sound trajectories in the embedding space, through the comparison with the state-of-the-art audio style transfers, as well as by the user studies. We expect that the novel music sounds synthesized with natural sounds by the proposed method can open new opportunities for audio-augmented and mixed reality (AR/MR).

Paper

Full text

PDF

Synthesis of Disparate Audio Species via Recurrent Neural Embedding

Semantic Scholar · Computer Science · 2023

Abstract

Sound synthesis plays a crucial role in the alteration of style, mood, and auditory affordances embedded in sounds. However, currently, sound synthesis has been limited mostly to a single species of sound, music sound. And, the synthesis of multiple disparate sound species that can transform the sound synthesis into another dimension is yet to be established. In this study, we present a method for synthesizing disparate sound species, in particular, music and natural sounds in real-time. We achieve this by interpolating the timbre embedding of music and natural sounds together with independent pitch conditioning, where a bidirectional LSTM stack is used to encode a full sequence of music and natural sounds while disentangling pitch from timbre. The effectiveness of the proposed sound synthesis is validated by evaluating the R2 score and the Structural Similarity Index Measure (SSIM) associated with synthetic sound trajectories in the embedding space, through the comparison with the state-of-the-art audio style transfers, as well as by the user studies. We expect that the novel music sounds synthesized with natural sounds by the proposed method can open new opportunities for audio-augmented and mixed reality (AR/MR).

Similar papers

© 2026 NYSGPT2525 LLC