On Exploring Audio Anomaly in Speech

Tiago Roxo, Joana Cabral Costa, Pedro R. M. Inácio, Hugo Proença · 2023

Existing anomaly detection works mainly focus on abnormal activities in image and video settings, while assessing audio manipulation, namely the presence of anomalous audio in speech, has not yet been explored. To overcome this limitation, we propose a setup in the context of Active Speaker Detection (ASD) by defining a methodology to perceive audio anomaly, assessing the performance of anomaly models, and establishing setup variations. This way, we evaluate models performance in identifying the presence of anomalies (detection) and localizing the timeframe where they occur (localization). To complement anomaly detection, we propose Anomaly Score (AS), a metric to assess anomaly localization that balances precision and mis-localization. Given the sequential nature of audio, we explore the performance of a density-based approach for video anomaly (CPD) and recurrent models (LSTM and RNN) on detecting and localizing audio anomalies. The results show that: 1) anomaly inclusion in talking portions increases models resilience toward anomaly localization; 2) CPD is superior in anomaly detection, while recurrent models perform better in anomaly localization; 3) anomaly with distinctive audio benefits precise anomaly localization; and 4) using original ASD audio is overall the best approach, relative to other processing approaches. The setup and experiments of this work serve as a baseline for future works on speech anomaly detection.

Read the paper · More papers on PaperTik