Attentional Multimodal Speech Enhancement for Voice User Interface in Consumer Electronics
Nasir Saleem, Sami Bourouis, Muhammad Irfan Khattak, Kia K. Dashtipour, Hela Elmannai, Ahmed Y. Al-Dubai, Tughrul Arslan, Amir Hussain · IEEE Transactions on Consumer Electronics · 2025
Voice User Interfaces (VUI) in consumer electronics often operate in noisy environments where speech quality degrades significantly. We propose AV-Net, a lightweight audiovisual speech enhancement model that addresses this challenge through novel cross-attentional feature fusion. Our approach dynamically integrates audio and visual modalities using a computationally efficient architecture combining a convolutional encoder-decoder (5.2M parameters) with a P3D-ResNet18 video encoder. The key innovation is a cross-attention mechanism that learns inter-modal correlations while preserving modality-specific features, outperforming conventional fusion methods. Evaluated on TCD-TIMIT and AVSE3 datasets under challenging conditions (SNR≤-5dB), AV-Net achieves a PESQ of 2.56 (vs. 1.26 baseline) and STOI improvement of 22%, while maintaining real-time performance (RTF=0.11). The model demonstrates strong generalization to unseen speakers and diverse noise types, making it particularly suitable for resource-constrained edge devices in healthcare, automotive, and smart home applications where robust speech interaction is critical.