Continuous sign language recognition based on cross-resolution knowledge distillation

The goal of continuous sign language recognition (CSLR) research is to apply CSLR models as a communication tool in real life, and the real-time requirement of the models is important. In this paper, we address the model real-time problem through cross-resolution knowledge distillation. We propose a new frame-level feature extractor that keeps the output frame-level features at the same scale as the output of by the teacher network. We further combined with the “TSCM + 2D” hybrid convolution proposed in our previous study to form a new lightweight end-to-end network-low-resolution input net (LRINet). It is then used to combine cross-resolution knowledge distillation to form the proposed CSLR model based on cross-resolution knowledge distillation (CRKD). The CRKD uses high-resolution frames as input to the teacher network for training, locks the weights after training, and then uses low-resolution frames as input to the student network LRINet to perform knowledge distillation on frame-level features and classification features, respectively. Experiments on two large-scale datasets have proved the effectiveness of CRKD. Compared to the ResNetT34 model with high resolution as input, it is 1.74 times faster than ResNetT34 in inference time and reduces the number of parameters and computation by 52.2% and 61.7%, while achieving comparable accuracy.

Paper

References (63)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC