Automatic Operating Room Surgical Activity Recognition for Robot-Assisted Surgery

Automatic recognition of surgical activities in the operating room (OR) is a\nkey technology for creating next generation intelligent surgical devices and\nworkflow monitoring/support systems. Such systems can potentially enhance\nefficiency in the OR, resulting in lower costs and improved care delivery to\nthe patients. In this paper, we investigate automatic surgical activity\nrecognition in robot-assisted operations. We collect the first large-scale\ndataset including 400 full-length multi-perspective videos from a variety of\nrobotic surgery cases captured using Time-of-Flight cameras. We densely\nannotate the videos with 10 most recognized and clinically relevant classes of\nactivities. Furthermore, we investigate state-of-the-art computer vision action\nrecognition techniques and adapt them for the OR environment and the dataset.\nFirst, we fine-tune the Inflated 3D ConvNet (I3D) for clip-level activity\nrecognition on our dataset and use it to extract features from the videos.\nThese features are then fed to a stack of 3 Temporal Gaussian Mixture layers\nwhich extracts context from neighboring clips, and eventually go through a Long\nShort Term Memory network to learn the order of activities in full-length\nvideos. We extensively assess the model and reach a peak performance of 88%\nmean Average Precision.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC