CRSS-LDNN: Long-duration naturalistic noise corpus containing multi-layer noise recordings for robust speech processing
John H. L. Hansen, Harishchandra Dubey, Abhijeet Sangwan · The Journal of the Acoustical Society of America · 2018
Multi-layer noise refers to scenarios where multiple distinct noise sources are simultaneously active in an audio stream. We collected a corpus named the CRSS long-duration naturalistic noise (CRSS-LDNN) corpus. It contains noise captured from complex daily-life activities using wearable LENA units. The diversity in noise-sources include construction noise, noise like multi-speaker babble, large-crowd noise, vehicle/bus noise on the road, home environment noise, etc. Corpus contains recordings of complex mixtures of these noise types. The babble noise and bus-engine noise present along with occasional impulsive-noise over long-duration is an example of such scenario. The data were collected at 16 kHz sampling rate with 16-bit precision in .wav format. This data would be released to speech community (http://crss.utdallas.edu). During the summer semester, a CRSS student wore a LENA device that was switched ON when multiple noise-sources were present. This corpus was recorded in naturalistic scenarios with uncontrolled mixing of various noise-sources. It provides naturalistic multi-layer noise recordings for evaluation of robust speech algorithms such as speech recognition, speaker diarization and verification, sentiment analysis. It consists of approximately 19 hours noise recordings. The CRSS-LDNN noise is more challenging as compared to existing noise corpora such as NOISEX that contains only single noise.