Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval

Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive captions. However, their effectiveness significantly degrades when handling lengthy textual descriptions, subjective queries, and complex acoustic environments, due to inherent architectural limitations and training data biases. To address these issues, we propose Nanami, a novel framework that capitalizes on multimodal large language models (MLLMs) to produce unified multimodal embedding through two synergistic mechanisms. First, Nanami aligns different modalities into a shared semantic space via designed prompting strategies and in-context learning (ICL) examples, thereby effectively bridging the modality gap. Second, Nanami refines through unimodal contrastive training in this space to enable efficient hybrid encoding, achieving strong retrieval performance. For comprehensive evaluation, we also introduce Nanamin, a new benchmark designed to assess model capabilities in processing long-form descriptions, subjectively expressed queries, and diverse acoustic samples. Experimental results confirm that Nanami, by utilizing the reasoning and generalization capabilities of MLLMs, achieves consistently robust performance in these complex tasks. The code and dataset will be available soon.

Paper

Full text

PDF

Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval

Semantic Scholar · Computer Science · 2025

Abstract

Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive captions. However, their effectiveness significantly degrades when handling lengthy textual descriptions, subjective queries, and complex acoustic environments, due to inherent architectural limitations and training data biases. To address these issues, we propose Nanami, a novel framework that capitalizes on multimodal large language models (MLLMs) to produce unified multimodal embedding through two synergistic mechanisms. First, Nanami aligns different modalities into a shared semantic space via designed prompting strategies and in-context learning (ICL) examples, thereby effectively bridging the modality gap. Second, Nanami refines through unimodal contrastive training in this space to enable efficient hybrid encoding, achieving strong retrieval performance. For comprehensive evaluation, we also introduce Nanamin, a new benchmark designed to assess model capabilities in processing long-form descriptions, subjectively expressed queries, and diverse acoustic samples. Experimental results confirm that Nanami, by utilizing the reasoning and generalization capabilities of MLLMs, achieves consistently robust performance in these complex tasks. The code and dataset will be available soon.

Similar papers

© 2026 NYSGPT2525 LLC