A Characteristic of Speaker's Audio in the Model Space Based on Adaptive Frequency Scaling
Jiamin Wu, Xiren Zhou, Qiuju Chen · 2022
The separation and extraction of the speaker's personality in the speech is one of the key links to determine the performance of the system in the specific research of speaker identification. In this paper, an adaptive frequency scaling-based identifiable feature in the model space is proposed which is suitable for speaker identification. Firstly, an adaptive filter is designed by introducing F -ratio to increase the frequency resolution of high-discrimination sub-bands based on the non-uniform distribution of speaker information in different frequency bands. The F -ratio is introduced to weight the filters and the feature called Non-Uniform Filter Bank coefficients (NUFBank) is obtained. After that, the Echo State Network (ESN) is used to learn the NUFBank feature sequence to obtain the state trajectory sequence. And then it is used as the representation of the speech signal in the model space for the distance measurement. The proposed characteristic has shown good interpretability and stability.