Gated Temporal Shifts with Depth-Efficient Channel Attention for Real-Time Hand-Gesture Interaction
Salah-eddine Laidoudi, Madjid Maïdi, Samir Otmane · 2025
We introduce a compact video-classification pipeline for real-time dynamic hand-gesture recognition in mixed-reality (MR) settings. The network marries a MobileNetV3 backbone with two purpose-built temporal components: (1) a Gated Discriminative Temporal Shift Module (G-DiTSM) that inserts first-order motion differences and learns channel-wise gates to fuse them adaptively, and (2) a lightweight Depth-Efficient Channel Attention (DepthECA) block that recalibrates spatial features on the fly. Operating on eight sparsely sampled frames per clip (Temporal Segment Network paradigm), the resulting model contains 2.65 M parameters and requires only 0.084 GFLOPs per inference. Evaluated on the RGB-only 20BN Jester benchmark (148k clips spanning 27 gesture classes) recorded from front-facing viewpoints. The system reaches 95.34% Top-1 and 99.80% Top-5 accuracy, surpassing recent 3D CNNs and transformer baselines while being an order of magnitude lighter. Ablations confirm that DepthECA and G-DiTSM provide complementary gains (+18.78% and +0.93% Top-1, respectively, over the MobileNetV3 baseline). Because all components are plug-and-play and introduce minimal overhead, the architecture is well suited to the tight latency and power budgets of standalone MR headsets, paving the way for natural grab, rotate, and command interactions using only on-board RGB cameras