MonoDepth-ViT: Enhancing Vision Transformer Robustness via Depth Spatial Features

Gusti Pangestu, Yaya Heryadi, Alexander Agung Santoso Gunawan, Widodo Budiharto · IEEE Access · 2025

Research in the domain of hand gesture recognition has undergone rapid expansion, with applications emerging in diverse fields, including User Experience (UX) and Human Computer Interaction (HCI). A prevalent approach is computer vision, which facilitates natural interaction without necessitating additional devices. In the domain of computer vision, variation of techniques exists for recognizing hand gestures. These techniques utilizes various approaches, including image classification and keypoint detection. However, it should be noted that keypoint-based methods are characterized by their high computational requirements. Consequently, the necessity for a more efficient mechanism with optimal accuracy has become apparent. A model that has been demonstrated to possess a high level of accuracy is the Vision Transformer (ViT). However, ViT tends to prioritize global context, which can limit its effectiveness when dealing with local context. Consequently, it struggles with images that contain high noise, such as images with varied backgrounds or those less centered on the object. The present study puts forth a series of modifications to ViT, with the objective of enhancing its capacity to recognize hand gestures, particularly in contexts characterized by intricate visual noise. The experimental results demonstrate that the proposed model exhibits enhanced robustness to noise and is capable of maintaining attention by focusing on the object, even in instances where the image is not focused on the object and possesses a highly variable background. The proposed model demonstrates the capacity to attain an average accuracy of 93.7%, thereby surpassing the 87% average accuracy achieved by the baseline model. This study makes novel contributions to the field of Vision Transformer architecture by demonstrating significant performance improvements through the integration of depth information into the input image processing pipeline.

Read the paper · More papers on PaperTik