Multi-Modal Mutual Information Maximization: A Novel Approach for Unsupervised Deep Cross-Modal Hashing
In this paper, we adopt the maximizing mutual information (MI) approach to\ntackle the problem of unsupervised learning of binary hash codes for efficient\ncross-modal retrieval. We proposed a novel method, dubbed Cross-Modal Info-Max\nHashing (CMIMH). First, to learn informative representations that can preserve\nboth intra- and inter-modal similarities, we leverage the recent advances in\nestimating variational lower-bound of MI to maximize the MI between the binary\nrepresentations and input features and between binary representations of\ndifferent modalities. By jointly maximizing these MIs under the assumption that\nthe binary representations are modelled by multivariate Bernoulli\ndistributions, we can learn binary representations, which can preserve both\nintra- and inter-modal similarities, effectively in a mini-batch manner with\ngradient descent. Furthermore, we find out that trying to minimize the modality\ngap by learning similar binary representations for the same instance from\ndifferent modalities could result in less informative representations. Hence,\nbalancing between reducing the modality gap and losing modality-private\ninformation is important for the cross-modal retrieval tasks. Quantitative\nevaluations on standard benchmark datasets demonstrate that the proposed method\nconsistently outperforms other state-of-the-art cross-modal retrieval methods.\n