Two-Stage Supervised Learning-Based Method to Detect Screams and Cries in Urban Environments

Anil Sharma, Sanjit K. Kaul · IEEE/ACM Transactions on Audio Speech and Language Processing · 2015

Smartphones can enable monitoring signs of distress as a human goes about his daily routine. Motivated by this possibility of 24x7 distress detection, we investigate detection of screaming and crying in urban environments, which we categorize into the contexts of indoors (home and office), outdoors, human conversation, large human gatherings, machinery, and audio from multimedia devices. Prior works are often restricted to specific environments or controlled settings. We propose a novel two-stage supervised learning based method, with tunable decision parameters for each stage, to achieve a desired true distress (scream and cry) detection rate (DR) and false alarm rate (FAR). We observe that the choice of the parameters is a function of the signal-to-noise ratio (SNR) of the distress signal, which is the ratio of the power of the distress signal to the power of the context audio. In the absence of SNR information, we show that a simple SNR estimation scheme performs well. Alternately, we show how the decision parameters can be selected based on the context estimated by the method. We show the results of testing the proposal over hundred hours of audio data recorded by the smartphones of ten volunteers as they went about their daily routines. Achieved performance is exemplified by a DR of 93.16% and a FAR of 4.76% at a SNR of 20 dB. The corresponding values for a SNR of 10 dB are 84.13% and 4.77%. Finally, we compare with, which also deals with audio event detection in noisy environments.

Read the paper · More papers on PaperTik