AI-Based Real-Time Sign Language to Text Converter Using Vision Transformer

People with speech and hearing impairments have communication difficulties that make it difficult for them to engage with society. Their main form of communication is sign language, which is difficult for outsiders to comprehend. There is a communication difficulty between people who has hearing impairments and that community continue to provide a lot of issue in daily encounters. In order to close this gap, the proposed system provide an AI-based real-time Sign Language to Text Converter that uses Machine Learning and a Vision Transformer (ViT) model. However, the lack of frame-level annotations co-articulation effects and temporal dependency modeling present difficulties for continuous sign streams. The Vision Transformer (ViT) with Spatial–Temporal Attention framework presented in this paper is intended for continuous streams of sign language without explicit segmentation. To automatically segment and convert sign sequences into natural language sentences the suggested architecture combines a spatial with the ViT encoder and by combining a Temporal Transformer attention with the ViT decoder and an alignment mechanism based on Connectionist Temporal Classification (CTC). According to experimental results the model outperforms conventional CNN–RNN pipelines in terms of translation accuracy and temporal coherence and generates sentence-level translations.

Paper

Full text

PDF

AI-Based Real-Time Sign Language to Text Converter Using Vision Transformer

Semantic Scholar · 2026

Abstract

People with speech and hearing impairments have communication difficulties that make it difficult for them to engage with society. Their main form of communication is sign language, which is difficult for outsiders to comprehend. There is a communication difficulty between people who has hearing impairments and that community continue to provide a lot of issue in daily encounters. In order to close this gap, the proposed system provide an AI-based real-time Sign Language to Text Converter that uses Machine Learning and a Vision Transformer (ViT) model. However, the lack of frame-level annotations co-articulation effects and temporal dependency modeling present difficulties for continuous sign streams. The Vision Transformer (ViT) with Spatial–Temporal Attention framework presented in this paper is intended for continuous streams of sign language without explicit segmentation. To automatically segment and convert sign sequences into natural language sentences the suggested architecture combines a spatial with the ViT encoder and by combining a Temporal Transformer attention with the ViT decoder and an alignment mechanism based on Connectionist Temporal Classification (CTC). According to experimental results the model outperforms conventional CNN–RNN pipelines in terms of translation accuracy and temporal coherence and generates sentence-level translations.

Similar papers

© 2026 NYSGPT2525 LLC