Deep Learning for Image to Sound Synthesis

Sound synthesis refers to the generation of sound using electronic hardware and software. Sound synthesis has many applications, such as creating electronic instruments that resemble real instruments or generate unusual, unique new sounds. Currently, these type synthesized instruments are commonly employed for music production of many different genres of music. The generation of a sound relies on selecting and adjusting many parameters related to the form of base waves, envelope signals, and digital sound filters. Unfortunately, the number of these parameters can be huge. Therefore, creating a new sound can be overwhelming, especially if the user is not familiar with the principles of wave formation. In this work, we present a method based on convolutional neural networks that can be used to generate synthesized sounds from images that depict geometrical figures. Our results indicate the feasibility of employing the proposed methodology for allowing users to create new sounds more easily.

Paper

Full text

PDF

Deep Learning for Image to Sound Synthesis

Semantic Scholar · Computer Science · 2020

Abstract

Sound synthesis refers to the generation of sound using electronic hardware and software. Sound synthesis has many applications, such as creating electronic instruments that resemble real instruments or generate unusual, unique new sounds. Currently, these type synthesized instruments are commonly employed for music production of many different genres of music. The generation of a sound relies on selecting and adjusting many parameters related to the form of base waves, envelope signals, and digital sound filters. Unfortunately, the number of these parameters can be huge. Therefore, creating a new sound can be overwhelming, especially if the user is not familiar with the principles of wave formation. In this work, we present a method based on convolutional neural networks that can be used to generate synthesized sounds from images that depict geometrical figures. Our results indicate the feasibility of employing the proposed methodology for allowing users to create new sounds more easily.

Similar papers

© 2026 NYSGPT2525 LLC