Cleaning, Segmenting, and Spell-Checking Text

Apress eBooks · 2009

W hen extracting text from different sources, you commonly end up with “noise” characters and unwanted whitespace. So you need tools to help you clean up this extracted text. For many applications, you’ll also want to segment text by identifying the boundaries of sentences and to spell-check text using a single suggestion or a list of suggestions. In this chapter, you’ll learn how to remove HTML tags, extract full text from an XML file, segment text into sentences, perform stemming and spell-checking, and recognize and remove noise characters. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik