A Hybrid Classification Approach For Artificial Speech Detection
Choon Beng Tan, Mohd Hanafi Ahmad Hijazi, Puteri N. E. Nohuddin · 2023
The emergence of voice biometrics has revolutionized user authentication methods by delivering enhanced security and convenience, steadily replacing less secure authentication techniques. However, the automatic speaker verification (ASV) systems remained vulnerable to spoofing attacks, especially artificial speech attacks. This is because artificial speech can be generated rapidly and in abundance using state-of-the-art speech synthesis and voice conversion algorithms. Recently, a one-class learning-based countermeasure known as the AIR-ASVspoof was introduced for artificial speech detection, which has proven to be the best-performing non-fusion system. As AIR-ASVspoof focuses solely on the bona fide class as the target in artificial speech detection, its practicality in a real-world environment may be restricted. The reason for this limitation lies in the fact that bona fide speech can be recorded in different environments, thereby increasing the risk of encountering false negatives. Hence, this paper proposes a hybrid approach to mitigate the issue, named Hybrid AIR-ASVspoof (HAIR-ASVspoof), by hybridizing end-to-end learning and classic machine learning. In this approach, deep features will be extracted from the dataset using the trained AIR-ASVspoof model and classified using a classic machine learning model such as Support Vector Machine (SVM). The objectives are: (i) to extract deep features using the AIR-ASVspoof; and (ii) to compare the performance of the AIR-ASVspoof and the proposed hybrid approach. An experiment was conducted to find the appropriate backend classifier for the proposed hybrid approach. Experimental results showed that the proposed approach outperformed the original AIR-ASVspoof, with an Equal Error Rate (EER) of 0.57%.