SILA: Signal-to-Language Augmentation for Enhanced Control in\n Text-to-Audio Generation

The field of text-to-audio generation has seen significant advancements, and\nyet the ability to finely control the acoustic characteristics of generated\naudio remains under-explored. In this paper, we introduce a novel yet simple\napproach to generate sound effects with control over key acoustic parameters\nsuch as loudness, pitch, reverb, fade, brightness, noise and duration, enabling\ncreative applications in sound design and content creation. These parameters\nextend beyond traditional Digital Signal Processing (DSP) techniques,\nincorporating learned representations that capture the subtleties of how sound\ncharacteristics can be shaped in context, enabling a richer and more nuanced\ncontrol over the generated audio. Our approach is model-agnostic and is based\non learning the disentanglement between audio semantics and its acoustic\nfeatures. Our approach not only enhances the versatility and expressiveness of\ntext-to-audio generation but also opens new avenues for creative audio\nproduction and sound design. Our objective and subjective evaluation results\ndemonstrate the effectiveness of our approach in producing high-quality,\ncustomizable audio outputs that align closely with user specifications.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC