Sign language recognition is a critical technology for enhancing communication accessibility for individuals with hearing impairments. In this paper, we present a robust and efficient system for sign language recognition using Long Short-Term Memory networks and Mediapipe, a versatile framework for machine learning solutions. Our approach leverages Mediapipe's pre-trained hand detection and tracking models to extract key hand landmarks from video sequences in real-time. These landmarks are then fed into an LSTM network, which is adept at capturing temporal dependencies, to classify various sign language gestures. There are learning aids available for those who are deaf or have trouble speaking or hearing, but they are rarely used. Real-time image processing would be used in the proposed system to handle live sign movements. After that, classifiers would be used to discriminate between different signs, and text would appear in the translated output. The existing systems can detect movements with a significant lag because they only use image processing. Our goal in our work is to develop a cognitive system that is trustworthy and sensitive enough for people with speech and hearing problems to use it in daily activities.
Paper
Full text
LSTM-Based Recognition of Sign Language
Semantic Scholar · 2024
Abstract
Sign language recognition is a critical technology for enhancing communication accessibility for individuals with hearing impairments. In this paper, we present a robust and efficient system for sign language recognition using Long Short-Term Memory networks and Mediapipe, a versatile framework for machine learning solutions. Our approach leverages Mediapipe's pre-trained hand detection and tracking models to extract key hand landmarks from video sequences in real-time. These landmarks are then fed into an LSTM network, which is adept at capturing temporal dependencies, to classify various sign language gestures. There are learning aids available for those who are deaf or have trouble speaking or hearing, but they are rarely used. Real-time image processing would be used in the proposed system to handle live sign movements. After that, classifiers would be used to discriminate between different signs, and text would appear in the translated output. The existing systems can detect movements with a significant lag because they only use image processing. Our goal in our work is to develop a cognitive system that is trustworthy and sensitive enough for people with speech and hearing problems to use it in daily activities.