Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuity
In robotic surgery, instrument presence labels are typically recorded alongside video streams, offering a cost-effective alternative to manual annotations for segmentation tasks. Label-supervised surgical instrument segmentation (SIS), a weakly supervised segmentation setting where only instrument presence labels are available, remains underexplored due to its inherently ill-posed nature. Temporal information plays a vital role in capturing sequential dependencies, thereby enhancing representation learning even under incomplete supervision. This paper extends a two-stage label-supervised segmentation framework by leveraging the temporal characteristics of surgical videos from three perspectives. First, a temporal equivariance constraint is introduced to enforce pixel-level consistency across adjacent frames. Second, a class-aware semantic continuity constraint is applied to preserve coherence between global and local regions over time. Third, temporally-enhanced pseudo masks are generated from consecutive frames to suppress irrelevant regions and improve segmentation accuracy. We evaluate our method on two surgical video datasets: the Cholec80 cholecystectomy benchmark and a real-world robotic left lateral segmentectomy (RLLS) dataset. Instance-level instrument annotations, sampled at regular intervals and validated by an experienced clinician, provide a reliable basis for evaluation. Experimental results demonstrate that our method consistently achieves favorable performances over state-of-the-art methods. These findings highlight the effectiveness of incorporating temporal constraints into label-supervised frameworks, offering a promising strategy to reduce annotation costs and advance surgical video analysis.