Multi-Vector Index Compression in Any Modality

We study efficient multi-vector retrieval for late interaction in any modality. Late interaction has emerged as a dominant paradigm for information retrieval across modalities, but its computation and storage costs grow linearly with document length, making it costly for multimodal corpora. To address this limitation, we explore methods for compressing multi-vector document representations under a constant vector budget. We introduce four approaches for index compression, including a novel attention-guided clustering (øurs). øurs uses an attention-guided mechanism to identify the most semantically salient parts of a document as cluster centroids and to weight token aggregation. Evaluating these methods on retrieval tasks spanning 4 modality settings (\beir, \vidore, \msrvtt, μltivent), we show that øurs consistently outperforms other parameterized compression methods (sequence resizing and memory tokens), provides greater flexibility in index size than non-parametric hierarchical clustering, and achieves competitive or improved performance compared to a full, uncompressed index.

Paper

Similar papers

© 2026 NYSGPT2525 LLC