Knowledge-Infused Retrieval Boosts Few-Shot Hand Gesture Recognition on HaGRID with Vision-Language Model
Enny Indasyah, Kaori Yoshida · Frontiers in artificial intelligence and applications · 2025
Hand gesture recognition remains challenging, primarily due to its reliance on large-scale annotated datasets and the limited adaptability of existing models when encountering novel gesture classes. In this work, we propose to apply Adaptive Vision-Language Model (Adaptive-VLM). This lightweight, training-free framework utilizes only one image per class to recognize gestures on the HaGRID benchmark. Built upon the CLIP backbone, our approach incorporates symbolic knowledge-infused prompts, multi-prompt contextualization, and semantic exemplar ranking to improve few-shot generalization. Adaptive-VLM achieves a macro F1-score of 65.75% on the HaGRID test set (540 images) without any parameter fine-tuning, using 18 example images. It significantly outperforms the Random-VLM baseline (59.95%) and a ResNet-18 model fine-tuned for 10 epochs (4.09%) under the same data constraints. These findings highlight the effectiveness of combining structured domain knowledge and guided exemplar selection to overcome data scarcity in low-resource gesture recognition. Adaptive-VLM offers a promising direction for building adaptive and efficient HGR systems, especially in real-world human-computer interaction scenarios requiring rapid deployment with minimal data.