Atom Aligned Sparse Representation Approach for Indonesian Emotional Speaker Recognition System
Andika Kusuma, Dessi Puji Lestari · 2020
Automatic Speaker Recognition system is a system that determines speaker identity through sound waves. This system can facilitate various daily services such as bank transaction via telephone. Nowadays, i-vector based Automatic Speaker Recognition system for Indonesian has not been able to handle the problem of emotional difference. However, in reality, speaker enrollment and recognition is often done in different emotional condition. This emotional difference frequently degrades the performance of existing systems. Therefore, this research focuses on constructing Automatic Speaker Recognition system for Indonesian that could handle emotional difference problem by applying i-vector modelling technique and Atom Aligned Sparse Representation (AASR) transformation technique. The emotion classes used in this study are angry, happiness, sadness, and contentment. Compared to the baseline system that was built using i-vector method only, the AASR system shows an increase in performance, namely a decrease in Equal Error Rate (EER) of 3.79% in non-neutral emotion test data. In neutral emotion test data, the AASR system also experiences a decrease in EER of 2.24%. Overall, the AASR system improves speaker recognition performance by reducing the EER by 3.46%.