Hybrid Instance-aware Temporal Fusion for Online Video Instance Segmentation

Recently, transformer-based image segmentation methods have achieved notable\nsuccess against previous solutions. While for video domains, how to effectively\nmodel temporal context with the attention of object instances across frames\nremains an open problem. In this paper, we propose an online video instance\nsegmentation framework with a novel instance-aware temporal fusion method. We\nfirst leverages the representation, i.e., a latent code in the global context\n(instance code) and CNN feature maps to represent instance- and pixel-level\nfeatures. Based on this representation, we introduce a cropping-free temporal\nfusion approach to model the temporal consistency between video frames.\nSpecifically, we encode global instance-specific information in the instance\ncode and build up inter-frame contextual fusion with hybrid attentions between\nthe instance codes and CNN feature maps. Inter-frame consistency between the\ninstance codes are further enforced with order constraints. By leveraging the\nlearned hybrid temporal consistency, we are able to directly retrieve and\nmaintain instance identities across frames, eliminating the complicated\nframe-wise instance matching in prior methods. Extensive experiments have been\nconducted on popular VIS datasets, i.e. Youtube-VIS-19/21. Our model achieves\nthe best performance among all online VIS methods. Notably, our model also\neclipses all offline methods when using the ResNet-50 backbone.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC