Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos

Ultrasound videos are an important form of clinical imaging data, and deep learning-based analysis can improve diagnostic accuracy and clinical efficiency. However, the scarcity of labeled data and the inherent challenges of video analysis have impeded the advancement of related methods. In this work, we introduce E-ViM3, a data-efficient Vision Mamba network that preserves the 3D structure of video data, enhancing long-range dependencies and inductive biases to better model spatial-temporal correlations. With our design of Enclosure Global Tokens (EG T), the model captures and aggregates global features more effectively than competing methods. We further employ a tailored masked video modeling approach for self-supervised pre-training to enhance its data efficiency, with the proposed Spatial- Temporal Chained (STC) masking strategy designed to adapt to different video scenarios. Experiments demonstrate that E-ViM3 achieves state-of-the-art performance on different tasks across four datasets of varying sizes: EchoNet-Dynamic, CAMUS, MICCAI-BUV, and WHBUS. Furthermore, our model attains competitive results even with limited labeled data, highlighting its potential impact on real-world clinical applications. Codes are available at https://github.com/HenryZhou19/E-ViM3.

Paper

Similar papers

© 2026 NYSGPT2525 LLC