Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!

Modeling expressive cross-modal interactions seems crucial in multimodal\ntasks, such as visual question answering. However, sometimes high-performing\nblack-box algorithms turn out to be mostly exploiting unimodal signals in the\ndata. We propose a new diagnostic tool, empirical multimodally-additive\nfunction projection (EMAP), for isolating whether or not cross-modal\ninteractions improve performance for a given model on a given task. This\nfunction projection modifies model predictions so that cross-modal interactions\nare eliminated, isolating the additive, unimodal structure. For seven\nimage+text classification tasks (on each of which we set new state-of-the-art\nbenchmarks), we find that, in many cases, removing cross-modal interactions\nresults in little to no performance degradation. Surprisingly, this holds even\nwhen expressive models, with capacity to consider interactions, otherwise\noutperform less expressive models; thus, performance improvements, even when\npresent, often cannot be attributed to consideration of cross-modal feature\ninteractions. We hence recommend that researchers in multimodal machine\nlearning report the performance not only of unimodal baselines, but also the\nEMAP of their best-performing model.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC