GlossFormer++: A Temporal Transformer for Sign Language Understanding Using 3D Hand Landmarks

Sign language recognition plays a vital role in enabling accessible communication for the Deaf and hard of hearing communities. While prior approaches using convolutional and recurrent models have shown progress, challenges remain in capturing fine grained hand dynamics and translating continuous gestures into semantically meaningful glosses. This work proposes GlossFormer++, a temporal transformer based architecture designed to recognize American Sign Language(ASL) glosses using 3D hand landmark sequences. This method processes 63-dimensional landmark vectors extracted via MediaPipe for each frame across temporally aligned sequences, enabling efficient modeling of motion and gesture transitions. To improve robustness and generalization, the model is pretrained on simulated ASL sequences from a high accuracy static hand sign classifier, followed by fine-tuning on the AUTSL gloss level dataset. Evaluated on normalized sequences of 32 frames, GlossFormer++ achieves a top 1 gloss classification accuracy of 96.8%, outperforming prior baselines using LSTMs and CNNs on the same landmark modality. This results highlight the potential of lightweight temporal transformers in real time, vision free sign language translation pipelines with low computational overhead and strong generalizability across signers.

Paper

Full text

PDF

GlossFormer++: A Temporal Transformer for Sign Language Understanding Using 3D Hand Landmarks

Semantic Scholar · 2025

Abstract

Sign language recognition plays a vital role in enabling accessible communication for the Deaf and hard of hearing communities. While prior approaches using convolutional and recurrent models have shown progress, challenges remain in capturing fine grained hand dynamics and translating continuous gestures into semantically meaningful glosses. This work proposes GlossFormer++, a temporal transformer based architecture designed to recognize American Sign Language(ASL) glosses using 3D hand landmark sequences. This method processes 63-dimensional landmark vectors extracted via MediaPipe for each frame across temporally aligned sequences, enabling efficient modeling of motion and gesture transitions. To improve robustness and generalization, the model is pretrained on simulated ASL sequences from a high accuracy static hand sign classifier, followed by fine-tuning on the AUTSL gloss level dataset. Evaluated on normalized sequences of 32 frames, GlossFormer++ achieves a top 1 gloss classification accuracy of 96.8%, outperforming prior baselines using LSTMs and CNNs on the same landmark modality. This results highlight the potential of lightweight temporal transformers in real time, vision free sign language translation pipelines with low computational overhead and strong generalizability across signers.

Similar papers

© 2026 NYSGPT2525 LLC