Research Objectives We propose a mobile application that recognizes particular programmed movements of the device to elicit particular text-to-speech phrases. For people unable to speak and with limited mobility to use keyboards or tablets, this enables faster social interaction which can greatly improve their quality of life. Design Initial prototype development and testing. Setting General community. Participants Six healthy subjects participated in repeated data collection of 11 distinct gestures for testing. Interventions Not applicable. Main Outcome Measures A cross-platform mobile application is integrated with a deep learning LSTM model that can recognize distinct gestures performed by the user after two seconds of recording, which triggers a selected text-to-speech auditory response. Results Eleven distinct gestures such as waving, horizontal line, vertical line, etc. data were collected through the accelerometer. Data were partitioned into three parts: training, validation and test. Training and validation data were used for training the model and the test set was used for evaluation. The model achieved 96% accuracy on average, with the errors primarily between one pair of challenging gestures to distinguish. Conclusions The application is an early, readily shared demonstration of gesture-to-speech. Further development is intended to enable additional gestures and improve recognition accuracy and speed, with the goal of enabling more efficient communication for individuals unable to speak. Author(s) Disclosures None.
Paper
Full text
A Gesture-to-Speech Recognition Mobile Application Prototype
Semantic Scholar · Computer Science · 2021
Abstract
Research Objectives We propose a mobile application that recognizes particular programmed movements of the device to elicit particular text-to-speech phrases. For people unable to speak and with limited mobility to use keyboards or tablets, this enables faster social interaction which can greatly improve their quality of life. Design Initial prototype development and testing. Setting General community. Participants Six healthy subjects participated in repeated data collection of 11 distinct gestures for testing. Interventions Not applicable. Main Outcome Measures A cross-platform mobile application is integrated with a deep learning LSTM model that can recognize distinct gestures performed by the user after two seconds of recording, which triggers a selected text-to-speech auditory response. Results Eleven distinct gestures such as waving, horizontal line, vertical line, etc. data were collected through the accelerometer. Data were partitioned into three parts: training, validation and test. Training and validation data were used for training the model and the test set was used for evaluation. The model achieved 96% accuracy on average, with the errors primarily between one pair of challenging gestures to distinguish. Conclusions The application is an early, readily shared demonstration of gesture-to-speech. Further development is intended to enable additional gestures and improve recognition accuracy and speed, with the goal of enabling more efficient communication for individuals unable to speak. Author(s) Disclosures None.