A Recurrent Neural Network-Based Multi-Parameter Sub-Band, Sub-Second Analysis for Voice Spoofing and Cloning Detection
A. Ramathilagam, R. Ramanathan, R. Saravanakumar · 2025
Speech synthesis, voice conversion, and replays are all forms of voice faking that might compromise Automatic Speaker Verification (ASV) systems. The ability to manipulate audio is a significant tool that criminals may use to take control of smart homes, bank accounts, and more. To detect Logical Access and Physical Access spoofing attacks using Multi-Pattern Features, this study employs a modified RNN architecture and Softmax layer. A robust voice spoofing detection system that can detect many types of spoofing attacks, intending to prevent ASV system fraud, has been provided. Local Binary Pattern Histogram (LBPH) is a novel feature descriptor that is provided for audio representation. By evaluating audio in both directions, LBPH can detect artificial speech artifacts, distortions in replay microphones, and dynamic speech characteristics in the real signal. Replay and logical-access attacks, such as speech synthesis and voice conversion, are detected by the Recurrent Neural Network (RNN) by using the recommended LBPH features. RNN is used since it is better at digesting and learning the representation of sequential input. It is discovered that physical-access attacks had an equal error rate of 0.69% and logical-access attacks with an accuracy of 90%. In addition to classifying speech cloning techniques, the proposed system can detect undiscovered voice spoofing attacks. When tested on the ASVspoof 2019 corpus, the suggested approach outperforms the state-of-the-art voice spoofing detection algorithms in detecting physical and logical access attacks.