Anomaly Detection in Voice Conversations Through Spectral Features

T Akilandeswari., N. Balaji, Paidimarri Nithish, Gosu Neeraj Yadav · 2024

Identifying anomalies in speech data presents a complex challenge due to various factors. Audio data is inherently high-dimensional and continuous. Background noise complicates the isolation of anomalies, while human perception of sound can be subjective. Speaker variability, including accents, dialects, and speaking speed, further complicates the task. The wide array of communication scenarios, ranging from informal conversations to professional settings, adds another layer of complexity. This research introduces a methodology that employs advanced signal processing techniques and machine learning algorithms to identify anomalies in audio communications. By utilizing features like Mel Frequency Cepstrum Coefficients (MFCC), spectral centroid, zero drift value, and power, the system extracts pertinent information from audio data to generate normal and abnormal speech patterns. Various statistical methods, including principal component analysis (PCA), Z-score, isolation forest, and local outlier factor (LOF), are employed to identify anomalous phenomena or behavior in voice interviews. The system offers a user-friendly interface for uploading audio files, visualizing extracted features, selecting anomaly detection methods, and interpreting detection results. Future enhancements encompass instant messaging integration, multi-specific vulnerability detection, and vulnerability scanning for police and cybersecurity agencies. It plays a vital role in ensuring the integrity, security, and quality of data transmission. The proposed model makes use of advanced unsupervised machine learning methods like Local Outlier Factor and Isolation Forest, which are adept at extracting complex, non-linear patterns from voice data without the need for labeled training data. This distinguishes it from more traditional techniques like HMMs, GMMs, and deep learning models, which may have trouble handling non-linear anomalies, need a large quantity of labeled data.

Read the paper · More papers on PaperTik