Spoof Speech Detection Based on Efficient Attention and Residual Network for Smart Consumer Electronics

Yongling Huang, Chenlu Zhu, Laurence Tianruo Yang, Yiheng Ruan, Xianjun Deng, Shenghao Liu, Rui Hou · IEEE Transactions on Consumer Electronics · 2025

In recent years, the development of technologies such as Artificial Intelligence and Internet of Things has provided richer interaction content for people’s interaction with smart consumer electronics. However, the traditional touchscreen interaction cannot meet the needs of people to achieve hands-free operation. Consequently, voice has emerged as a more widely used interaction medium. However, research demonstrates that intelligent voice systems are vulnerable to spoofing attacks, such as voice replay and synthetic voice attacks, which pose significant threats to user privacy and security. Existing spoof speech detection methods enhance multi-factor detection by analyzing additional signals such as sensor vibrations or Doppler shifts. However, these methods require specialized equipment or specific environmental conditions, making them difficult to deploy widely in complex scenarios. Furthermore, approaches relying solely on speech signals typically exhibit high computational complexity along with limited model robustness and generalization. This paper proposes a lightweight residual model based on efficient attention, which can effectively detect multiple types of spoof speech without causing extra burdens and can be applied to diverse smart consumer electronics. We orient the low-dimensional potential features extracted from residual networks for efficient attention computation in both channel and spatial dimensions to guarantee the model’s performance and efficiency. In addition, we comprehensively analyze various typical acoustic features to exploit their abilities to cope with spoof attacks. We conducted experiments on logical access (LA) and physical access (PA) datasets provided by the ASVspoof 2019 challenge. The Equal Error Rate (EER) is 43.02% lower than the baseline system on the LA dataset and 74.74% lower on the PA dataset.

Read the paper · More papers on PaperTik