Korean Dialect Identification Based on an Ensemble of Prosodic and Segmental Feature Learning for Forensic Speaker Profiling

Jooyoung Lee, Kyungwha Kim, Minhwa Chung · 2022

Telephonic fraud, or voice phishing, is becoming a major issue in South Korea, and there has been a high demand on AI-assisted solutions in the forensic areas to effectively narrow down the boundary of possible suspects. One such demand is use of automated Korean dialect identification in speaker profiling, which aims to classify dialect candidates of a suspect by analyzing dialectal patterns in speech, which are found in segmental and prosodic parts of the language. In this paper, an ensemble of dialect classifiers is proposed that considers both the segmental and the prosodic speech features. The classifier network is an ensemble of two sub-networks, which are an attention-based bidirectional LSTM network and a vanilla DNN for prosodic and segmental feature learning, respectively. A public dataset of Korean conversational speech is used, and the proposed model shows 61.28% in F-measure, which outperforms the baseline by 25.87%p.

Read the paper · More papers on PaperTik