Dual-Mode Robotic Arm Control System: Voice and Gesture Integration
Deepak Saha, N. K. Jisy, R Vinu, Selvam Sivasankari, Arun Balodi · 2025
Recent trends in robotics emphasize natural and efficient human-machine interaction, with a strong focus on multimodal systems that combine multiple input channels for enhanced control and adaptability. The main objective of this research is to contribute to the multimodal system by developing a cost-effective yet efficient solution that integrates both vision and voice modalities for robotic arm control, aiming to bridge the usability gap seen in traditional single-mode control systems. This research presents an innovative robotic control system that seamlessly integrates voice commands and hand gesture recognition to create an intuitive human-machine interface. This multimodal interaction framework synergistically integrates vision-based gesture recognition and voice command processing to overcome the inherent limitations of unimodal interfaces. The gesture subsystem leverages a convolutional neural network (CNN) to classify five hand poses with 86 % accuracy and$35 \mu$s latency, while the voice module employs the pretrained Vosk speech recognition model to achieve 94.8 % word accuracy with 120 ms mean response time. A temporal fusion algorithm synchronizes inputs within a 200 ms window, dynamically resolving conflicts through confidence-weighted arbitration. Empirical results demonstrate that the fused system reduces task completion time by 27 % in noisy environments compared to voice-only interfaces, while maintaining$\mathbf{9 1. 2 \%}$overall accuracy across 500 test samples.