Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation

Multiple object tracking and segmentation requires detecting, tracking, and\nsegmenting objects belonging to a set of given classes. Most approaches only\nexploit the temporal dimension to address the association problem, while\nrelying on single frame predictions for the segmentation mask itself. We\npropose Prototypical Cross-Attention Network (PCAN), capable of leveraging rich\nspatio-temporal information for online multiple object tracking and\nsegmentation. PCAN first distills a space-time memory into a set of prototypes\nand then employs cross-attention to retrieve rich information from the past\nframes. To segment each object, PCAN adopts a prototypical appearance module to\nlearn a set of contrastive foreground and background prototypes, which are then\npropagated over time. Extensive experiments demonstrate that PCAN outperforms\ncurrent video instance tracking and segmentation competition winners on both\nYoutube-VIS and BDD100K datasets, and shows efficacy to both one-stage and\ntwo-stage segmentation frameworks. Code and video resources are available at\nhttp://vis.xyz/pub/pcan.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC