Extraction of News Content for Text Mining Based on Edit Distance

Qiujun Lan · 2010

Online news as an up-to-date and important information source, is an absorbing data repository for data mining. However, news content of most web pages is embedded in a large amount of noisy materials. Accurate extraction of news content is a necessary and crucial step for news text mining. This paper proposes a new approach to news content extraction from web pages, which is based on several simple features observed in most well-known news websites/channels. One of the most important features is the similarity of the twin-pages which are collected from the same topic section of a site and published on the same/near date. A similarity measure based on edit distance is introduced and applied in the algorithms to separate the news content from noisy information. This method is much less complicated than other ones, and its accuracy and efficiency are fairly high, its complexity about the pages size is just linear. The experimental results based on 1514 pages collected from 54 top well-known Chinese and English news websites/channels show that it is very appropriate for most news pages to clean noise before text mining.

Read the paper · More papers on PaperTik