For individuals with hearing difficulties, Sign Language (SL) is an important communication tool. For non-signers, knowing is always an important obstacle. We describe an SL recognition model using the LSA64 dataset to tackle this problem. Our method utilizes six pre-trained models and a proprietary 3D Convolutional Neural Network (CNN) to capture spatial features. To efficiently capture temporal dependencies within video sequences of data, Features are further processed using a unified hybrid temporal modelling pipeline that incorporates a Temporal Convolutional Layer, a self-attention mechanism, and a Long Short Term Memory (LSTM) network. Robustness and comparability are preserved by using a uniform temporal architecture in all models. The suggested approach's 90.3 accuracy rate demonstrates promise and potency in automated Sign Language.
Paper
Full text
Sign Language Recognition from Video Using CNN Frame Feature Extraction and LSTM
Semantic Scholar · 2026
Abstract
For individuals with hearing difficulties, Sign Language (SL) is an important communication tool. For non-signers, knowing is always an important obstacle. We describe an SL recognition model using the LSA64 dataset to tackle this problem. Our method utilizes six pre-trained models and a proprietary 3D Convolutional Neural Network (CNN) to capture spatial features. To efficiently capture temporal dependencies within video sequences of data, Features are further processed using a unified hybrid temporal modelling pipeline that incorporates a Temporal Convolutional Layer, a self-attention mechanism, and a Long Short Term Memory (LSTM) network. Robustness and comparability are preserved by using a uniform temporal architecture in all models. The suggested approach's 90.3 accuracy rate demonstrates promise and potency in automated Sign Language.