Robust Voice Cloning Classification: Evaluating Monolingual and Regional Dialect Generalization
M Priyanka, G Rushita, G. Priyanka, P Amritha, Noor Aiman, R Suresha · 2025
Voice cloning technology has advanced significantly, posing a threat to the very basis of trust for voicebased systems, particularly in low-resource languages, where the reliable identification of synthetic speech is now highly crucial, as seen in languages like Kannada. Existing systems struggle to distinguish between artificial and real voices as voice cloning technology has become increasingly realistic. The proposed study presents a robust classification model that addresses the challenges of limited dataset size, phonetic variations, background noise, and environmental distortion. The model identifies major acoustic features such as MFCC, chroma, spectral contrast, and ZCR to absorb phonetic and tonal differences. Feature selection using LDA enhances class separability while improving accuracy and efficiency by eliminating redundant features. Among the classifiers tested, the SVM performed optimally with 98.2 % accuracy, 98.8 % precision, 97.3 % recall, and an F 1 -score of 98.0 %, better than other models on the Kannada regional dialect. The model was tested with existing datasets, such as Hindi and Malayalam, and demonstrated improved generalization. The study emphasizes the importance of enhanced voice authentication systems in aiding forensic investigations, ensuring inclusion, dependability, and flexibility to accommodate regional linguistic differences.