SAM-MPA: A SAM-Based Motion Perception and Aggregation Framework for Referring Video Segmentation

Referring video object segmentation relies on natural language descriptions to identify and segment target objects in videos, and has achieved substantial progress in recent years. However, most prior studies process videos in a frame-by-frame manner, failing to fully exploit temporal information. Recently, the large-scale segmentation model Segment Anything Model (SAM) has attracted considerable attention due to its strong segmentation capability and impressive zero-shot generalization. Nevertheless, SAM still exhibits limitations when handling complex action-oriented descriptions. Motivated by these observations, we propose a novel SAM-based Motion-Perception Aggregation framework for referring video object segmentation, termed SAM-MPA, which consists of four modules. DINO-SAM leverages the powerful segmentation ability of SAM to perform initial video segmentation guided by textual prompts, generating object-level masks. The Kalman Filtering Motion Modeling module injects explicit object motion modeling into DINO-SAM, improving segmentation robustness under occlusion and fast-motion scenarios. The motion-aware aggregation module effectively captures object action cues at multiple temporal scales, thereby enhancing global video understanding. The text-token matching module further enforces semantic consistency between the segmentation results and the referring expressions. Extensive experiments on challenging RVOS benchmarks demonstrate that SAM-MPA provides a competitive and efficient SAM-based solution for motion-centric referring video object segmentation, while offering a favorable trade-off between performance and computational cost compared with conventional non-MLLM baselines. The code is available at https://github.com/GXU-LIPE/SAM-MPA.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC