SYSTEM AND METHOD FOR VIDEO INSTANCE SEGMENTATION VIA RECURRENT ENCODER-BASED TRANSFORMERS
Patent №
US 12,657,917
Granted
2026-06-16
Filed 2024
Owner
Yeda Research and Development Co., Ltd
Lab
—
AI components
0
Assignment
None on record
Dataset
AIPD
Application
18435537
A method and system for video instance segmentation includes a recurrent encoder-based network trained by knowledge distillation from a transformer encoder. Real time performance is achieved by replacing the transformer encoder with the trained recurrent encoder for inference. The system includes a video camera to capture a sequence of video frames, a machine learning processing engine for video instance segmentation, and a video output for outputting a sequence of mask instances. The machine learning processing engine is configured with an interchangeable encoder module. During inference, the encoder module is configured with a recurrent encoder having a combination of convolutional and recurrent layers, The recurrent layers capture temporal relationships between the video frames. During training, the encoder module is configured with a teacher transformer encoder for training the recurrent encoder as a student through knowledge distillation. A transformer decoder outputs video instance mask predictions.
Ownership
Yeda Research and Development Co., Ltd