The nature of noise in linguistic corpora
Randy G. Goebel, Shane Bergsma, Ying Xu, Christoph Ringlstetter, Mi-Young Kim · 2010
'For the past 4-5 years, we've been investigating a variety of computational methods for extracting linguistic structures from relatively large language corpora. These include the use of well-known standard labeled language resources such as those from the Linguistic Data Consortium, as well as a spectrum of unlabeled resources, including the Google n-gram repository and a variety of more specific search engine query and answer resources (e.g., from Sogou).