Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies

Many recent datasets contain a variety of different data modalities, for\ninstance, image, question, and answer data in visual question answering (VQA).\nWhen training deep net classifiers on those multi-modal datasets, the\nmodalities get exploited at different scales, i.e., some modalities can more\neasily contribute to the classification results than others. This is suboptimal\nbecause the classifier is inherently biased towards a subset of the modalities.\nTo alleviate this shortcoming, we propose a novel regularization term based on\nthe functional entropy. Intuitively, this term encourages to balance the\ncontribution of each modality to the classification result. However,\nregularization with the functional entropy is challenging. To address this, we\ndevelop a method based on the log-Sobolev inequality, which bounds the\nfunctional entropy with the functional-Fisher-information. Intuitively, this\nmaximizes the amount of information that the modalities contribute. On the two\nchallenging multi-modal datasets VQA-CPv2 and SocialIQ, we obtain\nstate-of-the-art results while more uniformly exploiting the modalities. In\naddition, we demonstrate the efficacy of our method on Colored MNIST.\n

Paper

References (48)

Scroll for more · 36 remaining

Similar papers

© 2026 NYSGPT2525 LLC