Fearless steps: Taking the next step towards advanced speech technology for naturalistic audio

John H. L. Hansen, Aditya Joglekar, Abhijeet Sangwan, Chengzhu Yu · The Journal of the Acoustical Society of America · 2019

Over the past two decades, machine learning technologies have been targeting real-world problems in the Speech-Language (SLT) domain. Speech corpora developed under diverse environments have been paramount to progress, though most are simulated/controlled scenarios. Success in Machine Learning for SLT requires new innovative challenges posed by multi-speaker naturalistic audio. The UTDallas-CRSS led Fearless Steps (FS) initiative over the past 6 years has made significant strides through the development of the Fearless Steps Corpus, and multiple novel schemes to address core speech tasks. The next steps taken through this initiative aim at motivating further research on these data through worldwide collaborative efforts. To achieve this, the Inaugural Fearless Steps Challenge held in 2019 saw the release of 11 000 h of Apollo-11 audio data and diarization transcripts, made freely available to public online. An additional 100 h of manually annotated data was released as a Challenge corpus, available to researchers interested in working on any of the five core speech tasks: Speech-Activity-Detection, Speaker Diarization and Identification, Speech Recognition, and Sentiment Detection. FS Challenge resulted in over 150 participants worldwide developing novel task-specific algorithms. Going forward, the FS initiative aims at digitizing and releasing over 150 000 h of the remaining Apollo Missions in conjunction with follow-on Challenge Tasks focused on developing multi-channel and conversational speech systems.

Read the paper · More papers on PaperTik