People with speech and hearing impairments have communication difficulties that make it difficult for them to engage with society. Their main form of communication is sign language, which is difficult for outsiders to comprehend. There is a communication difficulty between people who has hearing impairments and that community continue to provide a lot of issue in daily encounters. In order to close this gap, the proposed system provide an AI-based real-time Sign Language to Text Converter that uses Machine Learning and a Vision Transformer (ViT) model. However, the lack of frame-level annotations co-articulation effects and temporal dependency modeling present difficulties for continuous sign streams. The Vision Transformer (ViT) with Spatial–Temporal Attention framework presented in this paper is intended for continuous streams of sign language without explicit segmentation. To automatically segment and convert sign sequences into natural language sentences the suggested architecture combines a spatial with the ViT encoder and by combining a Temporal Transformer attention with the ViT decoder and an alignment mechanism based on Connectionist Temporal Classification (CTC). According to experimental results the model outperforms conventional CNN–RNN pipelines in terms of translation accuracy and temporal coherence and generates sentence-level translations.
Paper
Full text
AI-Based Real-Time Sign Language to Text Converter Using Vision Transformer
Semantic Scholar · 2026
Abstract
People with speech and hearing impairments have communication difficulties that make it difficult for them to engage with society. Their main form of communication is sign language, which is difficult for outsiders to comprehend. There is a communication difficulty between people who has hearing impairments and that community continue to provide a lot of issue in daily encounters. In order to close this gap, the proposed system provide an AI-based real-time Sign Language to Text Converter that uses Machine Learning and a Vision Transformer (ViT) model. However, the lack of frame-level annotations co-articulation effects and temporal dependency modeling present difficulties for continuous sign streams. The Vision Transformer (ViT) with Spatial–Temporal Attention framework presented in this paper is intended for continuous streams of sign language without explicit segmentation. To automatically segment and convert sign sequences into natural language sentences the suggested architecture combines a spatial with the ViT encoder and by combining a Temporal Transformer attention with the ViT decoder and an alignment mechanism based on Connectionist Temporal Classification (CTC). According to experimental results the model outperforms conventional CNN–RNN pipelines in terms of translation accuracy and temporal coherence and generates sentence-level translations.