SYSTEM AND METHOD FOR VIDEO INSTANCE SEGMENTATION VIA RECURRENT ENCODER-BASED TRANSFORMERS

Patent №

US 12,657,917

Granted

2026-06-16

Filed 2024

Owner

Yeda Research and Development Co., Ltd

Lab

AI components

0

Assignment

None on record

Dataset

AIPD

Application

18435537

A method and system for video instance segmentation includes a recurrent encoder-based network trained by knowledge distillation from a transformer encoder. Real time performance is achieved by replacing the transformer encoder with the trained recurrent encoder for inference. The system includes a video camera to capture a sequence of video frames, a machine learning processing engine for video instance segmentation, and a video output for outputting a sequence of mask instances. The machine learning processing engine is configured with an interchangeable encoder module. During inference, the encoder module is configured with a recurrent encoder having a combination of convolutional and recurrent layers, The recurrent layers capture temporal relationships between the video frames. During training, the encoder module is configured with a teacher transformer encoder for training the recurrent encoder as a student through knowledge distillation. A transformer decoder outputs video instance mask predictions.

G06V 20/49G06V 10/774G06V 10/7715G06V 20/46G06V 10/82

Ownership

Yeda Research and Development Co., Ltd

© 2026 NYSGPT2525 LLC