AuraVox: mmWave-Augmented Audio Pipeline for Nonintrusive Emotion Sensing in Complex IoT Environments
Naveed Imran, Jian Zhang, Chihhsiong S. Shih, Sana Hameed, Abid Ishaq, Khursheed Aurangzeb · IEEE Internet of Things Journal · 2025
Reliable, privacy-preserving emotion sensing is essential for next-generation IoT applications; however, vision or audio-only pipelines often break down when faces are masked, environments are noisy, or multiple speakers coexist. We present AuraVox, a contactless system that couples millimetre-wave lip micro-Doppler with speech acoustics and fuses them through AuraNet, a bespoke cross-modal transformer. The radar first localizes each talker via MUSIC-based direction-of-arrival estimation and Bartlett beamforming, then captures high-resolution Doppler signatures of lip motion; in parallel, prosodic and spectral speech cues are extracted from a lapel microphone. AuraNet holistically attends to these heterogeneous streams and produces a unified representation for emotion classification. Evaluated on a 30-subject bilingual (English/Mandarin) corpus that includes mask-wearing, multi-speaker overlap, and 30–90 cm ranges, AuraVox attains 96% macro-F1, outperforming radar-only and audio-only baselines by up to eight percentage points. End-to-end latency is 12.8 ms per frame on a 10 W Jetson Xavier NX, meeting real-time constraints for edge deployment. By unifying beamformed lip kinematics with speech cues through AuraNet, AuraVox delivers the first multi-speaker, cross-language, mask-resilient emotion recognizer that runs on commodity hardware. Representative use cases include stress-aware in-cabin driver assistance, hospital check-in triage, and mood-adaptive smart-home interfaces.