Web Page Cleaning with Conditional Random Fields
Michal V. Marek, Pavel Pecina, Miroslav Spousta · 2007
This paper describes the participation of the Charles University in Cleaneval 2007, the shared task and competitive evaluation of automatic systems for cleaning arbitrary web pages with the goal of preparing web data for use as a corpus in the area of computational linguistics and natural language processing. We try to solve this task as a sequence-labeling problem and our experimental system is based on Conditional Random Fields exploiting a set of features extracted from textual content and HTML structure of analyzed web pages for each block of text. Labels assigned to these blocks then discriminate between different types of content blocks containing useful material that should be preserved in the document and noisy blocks of no linguistics interest that should be eliminated.