Pioneering Approaches for Accurate Vocal Anomaly Detection using State-of-the-Art Machine Learning Models

A Kannagi, Kalyan Acharjya, T. Prabhu · 2023

Accurate Vocal Anomaly Detection (AVAD) has emerged as a critical area of research with far-reaching implications for healthcare, security, and human-computer interaction. This paper presents pioneering approaches for AVAD using state-of-the-art machine learning models. Leveraging the power of Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and Support Vector Machines (SVMs), our proposed method demonstrates remarkable efficacy in identifying vocal anomalies indicative of various health conditions and security threats. In the data preprocessing phase, we extract Mel-Frequency Cepstral Coefficients (MFCCs) as vocal features, providing a robust foundation for subsequent analysis. The CNN, adept at capturing local patterns and hierarchies in feature maps, showcases impressive performance in classifying vocal data. LSTMs, renowned for capturing temporal dependencies in sequential data, enhance the model's ability to detect anomalies within vocal patterns. SVMs, with their capacity to construct optimal hyperplanes in feature space, serve as an additional valuable tool for accurate anomaly detection. Our comprehensive evaluation incorporates a range of performance metrics, including accuracy, precision, recall, F1 score, ROC-AUC, sensitivity, and specificity. The results demonstrate that our proposed method excels in accurately detecting vocal anomalies across diverse use cases, opening doors to early diagnosis, healthcare monitoring, and enhanced security measures through vocal analysis.

Read the paper · More papers on PaperTik