Fusion Strategies for Embedding Models: Enhancing Text Representations Across MTEB Tasks Through Lightweight and Trainable Ensembles

Text embeddings are critical components in modern NLP systems, underpinning tasks such as retrieval, classification, clustering, and semantic similarity. While individual embedding models have achieved notable success, their performance tends to vary across tasks due to differences in training data, objectives, and architectures. This work investigates whether combining multiple embedding models through ensemble fusion can produce more robust and general-purpose text representations.We benchmark five fusion strategies—averaging, concatenation, weighted averaging, MLP-based fusion, and trainable attention fusion—across over 15 tasks from the Massive Text Embedding Benchmark (MTEB). Our experiments show that even simple fusion techniques offer consistent performance gains over single-model baselines, while trainable fusion strategies, particularly attention-based fusion, deliver the highest overall performance.The study highlights the task-dependent behavior of each fusion method and reveals trade-offs between computational cost and representational strength. Our findings demonstrate that embedding-level ensembling is a practical, scalable, and effective approach to improving NLP performance across diverse benchmarks, ur findings demonstrate that embedding-level ensembling is a practical, scalable, and effective approach to improving NLP performance across diverse benchmarks. Quantitatively, trainable attention fusion achieves 9.8% higher nDCG@10 than averaging in retrieval tasks, while requiring only 2.1× more compute time than static methods. This work lays the groundwork for future adaptive fusion systems. laying the groundwork for future adaptive and multimodal fusion systems.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC