The NIST Meeting Room Pilot Corpus

John S. Garofolo, Christophe Laprun, Martial Michel, Vincent M. Stanford, Elham Tabassi · 2004

One of the next big challenges in Automatic Speech Recognition (ASR) is the transcription of speech in meetings.This task is particularly problematic for current recognition technologies because, in most realistic meeting scenarios, the vocabularies are unconstrained, the speech is spontaneous and often overlapping, and the microphones are inconspicuously placed.To support the development of meeting recognition technologies by both the speech recognition and video extraction research communities, NIST is providing a development and evaluation infrastructure including: a multi-media corpus of audio and video from meetings collected at NIST using a variety of microphones and video cameras, new evaluation protocols, metrics, software, rich transcription conventions, sponsoring evaluations and workshops, facilitating multi-site data pooling, and helping bring the community together to focus on the technical challenges.To date, NIST has collected a pilot corpus of 15 hours of meetings in its specially-instrumented Meeting Data Collection Laboratory.The corpus includes digital recordings from close-talking mics, lapel mics, distantly-placed mics, 5 digitally-recorded camera views, and full speaker/word-level transcripts.This data is being used in the development and evaluation of speech technologies and by the video extraction community under the auspices of the ARDA Video Analysis and Content Exploitation (VACE) program. MotivationHuge efforts are being expended in mining information in newswire, news broadcasts, and conversational speech and in developing interfaces to metadata extracted in these domains.However, little has been done to address such applications in the more challenging and equally important meeting domain.The development of smart meeting room core technologies that can automatically recognize and extract important information from multi-media sensor inputs will provide an invaluable resource for a variety of business, academic, and governmental applications.Such metadata will provide the basis for the development of second-tier meeting applications that can automatically process, categorize, and index meetings.Third-tier applications will provide a context-aware collaborative interface between live meeting participants, remote participants, meeting archives and vast online resources.Given that the necessary core meeting recognition technologies are in a fledgling or nonexistent state, it is essential that these first tier technologies be developed before the higher tier applications can be made useful.The meeting domain has several important properties not found in other domains and which are not currently being focused on in other research programs:• Multiple Forums and Vocabularies --Meeting forums range from very informal to highly structured.Likewise, meeting vocabularies vary widely depending on both the meeting topic and degree of shared context among the participants.• Highly-Interactive/Simultaneous Speech --The speech found in informal meetings is spontaneous and highly interactive across multiple participants and contains frequent interruptions and overlapping speech.This poses great challenges to speech recognition technologies that are typically tailored for singlespeaker speech streams.• Multiple Distant Microphones -Meetings are typically recorded with multiple distant microphones.Speech recognition systems generally work quite poorly with distant mics.Moreover, techniques have yet to be developed which efficiently integrate input from multiple mics and take advantage of their positioning to improve recognition quality.• Multiple Camera Views -Meetings are often recorded with multiple cameras with different and sometimes overlapping views.Much like the multi-microphone challenge above, this permits/challenges the technology to integrate data from multiple video inputs to enhance the metadata that can be extracted from the meeting and improve recognition quality.• Multi-Media Information Integration --It is impossible to develop a complete understanding of meetings without analyzing a number of different signal types simultaneously: audio, video and other information sources (devices/resources participants interact with).

Read the paper · More papers on PaperTik