An Intelligent Vision Transformer and Attention-Based Framework Using GuVAST-Net for Gujarati Optical Character Recognition and Speech Synthesis

Textual information, such as notice boards, notifications, learning resources, and visualisations of general information, can be difficult for people with visual impairment to read and interpret. While the deployment of text-to-speech (TTS) and optical character recognition (OCR) systems is helping to make data more accessible, the regulations of real-world scenarios, including variances of light intensities, strength of text backgrounds, font patterns, and the layout of characters, make the reliable recognition of Gujarati text in real scenarios challenging. Accordingly, the technological nature of a high-level assistance framework that distinguishes Gujarati speech and conveys information is a pressing requirement for two wearable device access programs. The proposed GuVAST-Net framework includes image capture, pre-processing, adaptive text segmentation, BiLSTM–Attention-based sequence modelling, Hybrid CNN–Vision Transformer (CNN–ViT) feature extraction, and Bayesian Optimisation–Reinforcement Learning (BO–RL) driven optimisation. Vision Transformer captures global contextual relations, whereas CNN restores local Gujarati characterisation representations. The second functionality of this comprises identifying Gujarati text in the speech, translating the identified text into recognisable spoken language for users with eyesight control, and, keeping in mind the asymmetric nature of speech translation, the presence of BiLSTM and Attention yields better performance in sequential recognition. For experimental evaluation, a dataset of Gujarati character classes was selected, with sample counts ranging from 115 to 375 per class. The proposed design achieved 96% recognition accuracy, 98% training accuracy, and 96% verification accuracy. Bayesian Optimisation reduces the loss from 0.99 to 0.11, an 88.9% improvement, while the standard training loss falls from 0.86 to 0.05, a 94.2% reduction. The RL agent demonstrated efficient minimisation and educational consistency, increasing the cumulative reward from −5 to 10 across a series of tests. The outcomes reveal that the proposed Hybrid CNN–ViT–BiLSTM architecture provides a reliable Voice-Based assistive system and respectable Gujarati text recognition, which is appropriate for real-time adaptive glasses for blind individuals.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC