MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization
In this work we present a new State-of-The-Art on the text-to-video retrieval\ntask on MSR-VTT, LSMDC, MSVD, YouCook2 and TGIF obtained by a single model.\nThree different data sources are combined: weakly-supervised videos,\ncrowd-labeled text-image pairs and text-video pairs. A careful analysis of\navailable pre-trained networks helps to choose the best prior-knowledge ones.\nWe introduce three-stage training procedure that provides high transfer\nknowledge efficiency and allows to use noisy datasets during training without\nprior knowledge degradation. Additionally, double positional encoding is used\nfor better fusion of different modalities and a simple method for non-square\ninputs processing is suggested.\n