M-Diarization: A Myanmar Speaker Diarization using Multi-scale dynamic weights

Myat Aye Aye Aung, Win Pa Pa, Hay Mar Soe Naing · 2023

This paper introduces M-Diarization dataset for Myanmar language. The dataset contains the dialog of two speakers in a meeting, live interview, discussion and talk show. The primary objective of this dataset is to facilitate speaker diarization, speech/speaker recognition, as well as speaker identification and verification tasks. Furthermore, it can serve as a valuable resource for various tasks, including but not limited to dereverberation, denoising, and speech enhancement, supporting a wide range of Myanmar speech processing research. The dataset is applied in the study for Myanmar speaker diarization. A trade-off in speaker diarization systems is struggled to achieve both precise speaker turn boundaries and accurate speaker identity information. Short speech segments and overlapped speech that are challenging for machines to transcribe and identify the speaker. A multi-scale strategy addresses this trade-off by extracting speaker features from segments of various lengths and subsequently integrating the outcomes from these multiple scales. In this paper, applied three different sets of multi-scale weight values, each with varying lengths (5, 6, and 7 scales), to analyze diverse M-Diarization Datasets and utilized a pre-trained speaker embedding extractor, specifically TitaNet-L, to predict the number of speakers and derive average speaker representation vectors for each speaker at every scale. Each scales length includes two types of parameters. The contribution in this study that present new speaker diarization dataset and significantly improve the diarization performance as a baseline on different length and multi-scale weight values. There are six parameters P1 (baseline), P2, P3, P4, P5 and P6. Parameter P2 of base scale 0.5s and shift length 0.25s is outperformed than the other types of multi-scale weight values. This system achieves 6.64% Diarization Error Rate (DER) on 3 hours, 7.89% DER on 5 hours and 9.33% DER on 8 hours on M-Diarization dataset.

Read the paper · More papers on PaperTik