The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment

Bird-MML is a synthetic multimodal dataset designed to study cross-modal representation learning and multimodal complementarity across vision, audio, and text. Despite substantial progress in multimodal learning, there remains a lack of standardized datasets that support the evaluation of both pairwise and higher-order multimodal alignment under controlled and scalable conditions. Bird species provide a natural domain for multimodal learning because their identity and behavior are expressed through both visual characteristics (e.g., morphology and plumage) and acoustic signals (e.g., vocalizations and calls). In many cases, fine-grained recognition depends on complementary information across modalities, making the domain well suited for studying multimodal interaction. To facilitate such research, we construct Bird-MML, a dataset containing artificially aligned image–audio–text triplets for 150 bird species, including all species present in the SSW60 and VB100 benchmarks. The dataset is designed for large-scale multimodal pretraining and cross-modal representation learning, with balanced coverage across species. Dataset Composition The dataset contains 149,681 multimodal triplets, each consisting of: Image – a bird photograph Audio – a 10-second bird vocalization clip Text – a multimodal caption describing the sample Each sample is associated with a species identity, enabling supervised or self-supervised multimodal learning tasks. The dataset maintains approximately 1,000 samples per species to ensure balanced representation. Data Sources Bird-MML integrates data from multiple publicly available sources: Images Images are sourced from the iNaturalist open dataset, restricted to research-grade observations released under Creative Commons licenses. Audio Audio recordings are obtained from Xeno-Canto, segmented into 10-second clips and zero-padded when necessary. When recordings were limited for a species, clips were reused to maintain balanced sample counts. Text Descriptions Text captions are generated by combining three complementary sources: Image captions generated by InstructBLIP2 Audio metadata (e.g., call type, sex, life stage) Short species summaries derived from Wikipedia These elements are fused using google/gemma-2-2b-it to produce a single multimodal caption describing each sample. Dataset Characteristics Bird-MML contains synthetically aligned multimodal triplets, meaning that the image, audio, and text components are algorithmically paired rather than originating from the same real-world observation. Although generated captions may contain minor factual noise, they capture species-relevant attributes and contextual cues that are useful for multimodal representation learning. The dataset is particularly suited for research on: multimodal representation learning contrastive learning across modalities cross-modal retrieval multimodal alignment methods Models trained using Bird-MML can be evaluated on naturally paired audio-visual datasets to assess real-world cross-modal generalization. Licensing and Attribution Bird-MML is constructed from publicly available datasets and resources released under Creative Commons licenses. Each individual sample retains the license of its original source. The creators of Bird-MML do not claim ownership of the underlying images, audio recordings, or textual sources. Ownership and attribution remain with the original contributors, including iNaturalist observers, Xeno-Canto recordists, and Wikipedia contributors. Source attribution and license information are preserved in the dataset metadata. Users of the dataset are responsible for respecting the original licenses associated with each sample, including attribution requirements and any restrictions on commercial use. Bird-MML is released for non-commercial research purposes only. Ethical Considerations Automatically generated captions may contain minor inaccuracies or biases and should not be interpreted as authoritative biological information. The dataset is intended exclusively for machine learning and multimodal representation learning research. Citation If you use the Bird-MML dataset in your research, please cite the following paper: @article{koutoupis2025more, title={The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment}, author={Koutoupis, Stefanos and Zervou, Michaela Areti and Kontras, Konstantinos and De Vos, Maarten and Tsakalides, Panagiotis and Tsagatakis, Grigorios}, journal={arXiv preprint arXiv:2511.21331}, year={2025} }

Paper

Similar papers

© 2026 NYSGPT2525 LLC