Speaker Identification for Detection of Mild Cognitive Impairment Using Bilingual Voice Recordings of Picture Description Tasks

Hsiao-Lun Wang, Hong-Han Hank Chau, Yawgeng A. Chau · 2024

For the detection of mild cognitive impairment (MCI), we study speaker identification based on unsupervised learning with clustering. The voice of potential patients in picture description tasks was analyzed in Mandarin or English. The dataset included 67 Taiwanese and 62 English recordings. The speaker embeddings from Pyannote were used to identify the speaker type. We used three different clustering methods, namely K-means, K-medoid, and agglomerative clustering. To examine the performance of the clustering methods, we evaluated and compared the metrics of the silhouette score and the Rand index. Based on the evaluation of the Rand index, agglomerative clustering yielded the highest performance score for the clustering of speaker embeddings.

Read the paper · More papers on PaperTik