Lip Reading Using Convolutional Auto Encoders as Feature Extractor

Visual recognition of speech using the lip movement is called Lip-reading. Recent developments in this nascent field use different neural networks as feature extractors which serve as input to a model which can map the temporal relationship and classify. Though an end to end sentence level Lip-reading is the current trend, we proposed a new model which employs word level classification and breaks the set benchmarks for standard datasets. In our model, we use convolutional autoencoders as feature extractors which are then fed to a Long short-term memory model. We tested our proposed model on BBC’s LRW dataset, MIRACL-VC1, and GRID dataset. Achieving a classification accuracy of 98% on MIRACL-VC1 as compared to 93.4% of the set benchmark by Rekik et al. On BBC’s LRW the proposed model performed better than the baseline model of convolutional neural networks and Long short-term memory model as seen in Garg et al. Showing the features learned by the models we clearly indicate how the proposed model works better than the baseline model. The same model can also be extended for an end to end sentence level classification.

Paper

Similar papers

© 2026 NYSGPT2525 LLC