Augmenting Vision-Language Retrieval: The Role of Multimodal LLMs as Synthetic Data Generators

Multimodal Large Language Models (MLLMs) connect and interpret different data types, making them suitable for various vision-language tasks. Despite the rapid advancements in MLLMs, their effectiveness for specialized cross-modal retrieval tasks remains underexplored. A challenging example is art retrieval, where the task is to find visually and conceptually relevant artwork corresponding to a textual description. This paper investigates the effects of fine-tuning cross-modal retrieval models using both human-annotated and MLLM-generated captions for artistic paintings. To this end, two cross-modal retrieval models, Long-CLIP and BLIP, are studied. Experimental results show that models fine-tuned on MLLM-generated captions achieve search effectiveness comparable to those fine-tuned on human-annotated captions.

Paper

Full text

PDF

Augmenting Vision-Language Retrieval: The Role of Multimodal LLMs as Synthetic Data Generators

OpenAlex · Multimodal Machine Learning Applications · 2025

Abstract

Multimodal Large Language Models (MLLMs) connect and interpret different data types, making them suitable for various vision-language tasks. Despite the rapid advancements in MLLMs, their effectiveness for specialized cross-modal retrieval tasks remains underexplored. A challenging example is art retrieval, where the task is to find visually and conceptually relevant artwork corresponding to a textual description. This paper investigates the effects of fine-tuning cross-modal retrieval models using both human-annotated and MLLM-generated captions for artistic paintings. To this end, two cross-modal retrieval models, Long-CLIP and BLIP, are studied. Experimental results show that models fine-tuned on MLLM-generated captions achieve search effectiveness comparable to those fine-tuned on human-annotated captions.

References (19)

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC