Multi-site data collection for a spoken language corpus - MAD COW

Lynette Hirschman · 1992

This paper describes a recently collected spoken language corpus for the ATIS (Air Travel Information System) domain. This data collection effort has been co-ordinated by MADCOW (Multi-site ATIS Data COllection Working group). We summarize the motivation for this effort, the goals, the implementation of a multi-site data collection paradigm, and the accomplishments of MADCOW in monitoring the collection and distribution of 12,000 utterances of spontaneous speech from five sites for use in a multi-site common evaluation of speech, natural language and spoken language. 1. Introduction Following the February 1991 DARPA Speech and Natural Language Workshop, the DARPA Spoken Language contractors decided to institute a multi-site data collection paradigm in order to: ffl support a common evaluation on speech, natural language and spoken language; ffl maximize the amount of data collected; ffl provide some diversity in data collection paradigms; ffl reduce cost to any one site by sharing...

Read the paper · More papers on PaperTik