This paper presents a system that combines hand gesture and audio recognition using deep learning techniques. Convolutional Neural Networks (CNN) are used by the system to recognise hand movements, while an architecture based on CNN + RNN and the CTC Loss algorithm is used to recognise audio. The proposed system is capable of recognizing sign language gestures and individual characters in speech, which can be helpful for individuals with hearing and speech impairments. The system is trained and tested on publicly available datasets, achieving high accuracy rates in both gesture and audio recognition. This work could advance the field of human-computer interaction and increase accessibility for people with hearing or speech problems. The outcomes show how well the suggested method works and illustrate its potential for use in practical situations. The results demonstrate that our model reliably recognises audio signals at a state-of-the-art level when tested against a public dataset.
Paper
Full text
Hand Gesture and Audio Recognition System Using Neural Networks
Semantic Scholar · Computer Science · 2023
Abstract
This paper presents a system that combines hand gesture and audio recognition using deep learning techniques. Convolutional Neural Networks (CNN) are used by the system to recognise hand movements, while an architecture based on CNN + RNN and the CTC Loss algorithm is used to recognise audio. The proposed system is capable of recognizing sign language gestures and individual characters in speech, which can be helpful for individuals with hearing and speech impairments. The system is trained and tested on publicly available datasets, achieving high accuracy rates in both gesture and audio recognition. This work could advance the field of human-computer interaction and increase accessibility for people with hearing or speech problems. The outcomes show how well the suggested method works and illustrate its potential for use in practical situations. The results demonstrate that our model reliably recognises audio signals at a state-of-the-art level when tested against a public dataset.