Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive captions. However, their effectiveness significantly degrades when handling lengthy textual descriptions, subjective queries, and complex acoustic environments, due to inherent architectural limitations and training data biases. To address these issues, we propose Nanami, a novel framework that capitalizes on multimodal large language models (MLLMs) to produce unified multimodal embedding through two synergistic mechanisms. First, Nanami aligns different modalities into a shared semantic space via designed prompting strategies and in-context learning (ICL) examples, thereby effectively bridging the modality gap. Second, Nanami refines through unimodal contrastive training in this space to enable efficient hybrid encoding, achieving strong retrieval performance. For comprehensive evaluation, we also introduce Nanamin, a new benchmark designed to assess model capabilities in processing long-form descriptions, subjectively expressed queries, and diverse acoustic samples. Experimental results confirm that Nanami, by utilizing the reasoning and generalization capabilities of MLLMs, achieves consistently robust performance in these complex tasks. The code and dataset will be available soon.
Paper
Full text
Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval
Semantic Scholar · Computer Science · 2025
Abstract
Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive captions. However, their effectiveness significantly degrades when handling lengthy textual descriptions, subjective queries, and complex acoustic environments, due to inherent architectural limitations and training data biases. To address these issues, we propose Nanami, a novel framework that capitalizes on multimodal large language models (MLLMs) to produce unified multimodal embedding through two synergistic mechanisms. First, Nanami aligns different modalities into a shared semantic space via designed prompting strategies and in-context learning (ICL) examples, thereby effectively bridging the modality gap. Second, Nanami refines through unimodal contrastive training in this space to enable efficient hybrid encoding, achieving strong retrieval performance. For comprehensive evaluation, we also introduce Nanamin, a new benchmark designed to assess model capabilities in processing long-form descriptions, subjectively expressed queries, and diverse acoustic samples. Experimental results confirm that Nanami, by utilizing the reasoning and generalization capabilities of MLLMs, achieves consistently robust performance in these complex tasks. The code and dataset will be available soon.