SignViT: An enhanced vision transformer framework for Attention-Based sign language hand gesture recognition
Umar Ashfaq, Qingshan Wang, Badreddine Merabet, Jiangtao Zhang · Biomedical Signal Processing and Control · 2025
Hand gesture recognition is an important part of the automation of sign languages. Capturing long-range dependencies (LRDs) in image recognition tasks is ill-suited in existing models to recognize the visually similar hand gesture images with high recognition accuracy. This paper proposes SignViT, a model best-suited for capturing LRDs, which has a strong backbone based on the vision transformer (ViT) to address the issue of visual similarity. Processing the images' features through the attention mechanism of the transformer's encoder block and a novel enhanced feedforward neural network classification head in our framework makes it highly accurate for recognizing all hand gestures, including visually similar gestures. To evaluate the model's performance, we assessed it on four hand gesture datasets representing three distinct sign languages, including Urdu, American, and Arabic. In our evaluation, we find that SignViT achieves state-of-the-art results on both similar-looking gestures and overall datasets compared to baseline ViT, convolutional neural network approaches, and other models. The overall accuracies achieved by SignViT on four datasets are 99.70%, 99.44%, 99.90%, and 99.52%.