Cross-Domain Few-Shot Hand Gesture Recognition on Edge Devices via Lightweight CLIP-KNN Fusion
Zenghui Yin · Applied and Computational Engineering · 2025
Hand-gesture interfaces are rapidly proliferating in AR/VR headsets, in-vehicle infotainment and battery-powered mobile devices. Yet large-scale annotated datasets are expensive to collect on edge hardware due to privacy, latency and energy constraints. We present CLIP-KNN Fusion, a parameter-efficient 5-shot adapter that achieves 93.2 percent top-1 accuracy at 32 FPS on an RTX 3050 4 GB laptop under unseen lighting, users and backgrounds. The system keeps the 86 M-parameter ViT-B/32 CLIP vision encoder completely frozen, and trains only a 5 kB cosine-distance 3-nearest-neighbor classifier on 25 RGB images (5-way × 5-shot). Compared to a CNN trained from scratch, we improve accuracy by 25.2 % while reducing trainable parameters by six orders of magnitude. Extensive cross-domain experiments confirm robustness to illumination shifts, novel users and cluttered backgrounds, echoing recent findings in cross-domain object detection. The plug-and-play codebase, 25-shot dataset and live webcam demo are released anonymously. To our knowledge, this is the first work to combine a frozen CLIP encoder with an on-device KNN head for real-time gesture recognition at the extreme edge.