Robust processing of noisy web-collected data

Jelke Bloem, Michaela Regneri, Stefan Thater · 2012

Crowdsourcing has become an important means for collecting linguistic data. However, the output of web-based experiments is often challenging in terms of spelling, grammar and out-of-dictionary words, and is therefore hard to process with standard NLP tools. Instead of the common practice of discarding data outliers that seem unsuitable for further processing, we introduce an approach that tunes NLP tools such that they can reliably clean and process noisy data collected for a narrow but unknown domain. We demonstrate this by modifying a spell-checker and building a coreference resolution tool to process data for paraphrasing and script learning, and we reach state-of-the-art performance where the original state-of-the-art tools fail.

Read the paper · More papers on PaperTik