THE TDT-3 TEXT AND SPEECH CORPUS

David Graff, Chris Cieri, Stephanie M. Strassel, Nii Martey · 2007

The TDT-3 Text and Speech Corpus expands on previous phases of Topic Detection and Tracking data collections, by increasing the number of news sources being sampled, by including Mandarin Chinese as well as English news data, and by introducing new forms of topic annotation. In order to satisfy the specific data and annotation requirements of the TDT-3 Evaluation Plan[1], the LDC refined and supplemented the methods that had been used in TDT-2 corpus development[2]. There were significant changes and improvements in the process of selecting anddefining target topics,in the procedures for quality assurance applied to both data content and annotations, and in the organization of the delivered corpus. In addition, the LDC created or acquired a range of resources to support research in cross-language information retrieval. These included the addition of a Mandarin Chinese component to the TDT-2 Text and Speech Corpus, the collection of a large body of Chinese-English parallel text, and ada...

Read the paper · More papers on PaperTik