Robust Scream Detection and Scream Temporal Interval Prediction Using CNN-Transformer and Windowing CNN
Youngjun Kim, Dalwon Jang, Seok-Pil Lee · IEEE Access · 2025
Detecting screams in noisy environments remains an important challenge in emergency detection. This work presents a new hybrid architecture combining Convolutional Neural Networks (CNN)-Transformer and Windowing CNN for simultaneous scream detection and temporal localization. The proposed system employs parallel CNN blocks for robust feature extraction alongside a Transformer block that captures temporal dependencies in the audio signal. A specialized Windowing CNN with Huber Loss regression enables precise temporal predictions of scream intervals. Experimental results demonstrate significant advancements over existing approaches, achieving approximately 3 percentage point improvement in F-measure and a remarkable 10-fold enhancement in Equal Error Rate (EER). To enhance real-world applicability, we developed an extensive dataset with various noise types at equal Signal-to-Noise Ratio (SNR), incorporating Time Shifting augmentation techniques. The model was evaluated across diverse acoustic scenarios including crowd, factory, white, and road noise environments. The proposed architecture effectively addresses both detection accuracy and temporal localization challenges under real-world noise conditions, offering substantial reliability improvements for practical security applications.