DOM-based print-link detection for web article extraction

Sam Liu, SukHwan Lim, Jerry Liu · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2011

Web article pages usually have hyperlinks (or links) that lead to print-friendly web pages containing mainly the article content. Content extraction using these print-friendly pages is generally easier and more reliable, but there are many variations of the print-link representations in HTML that made robust print-link detection more difficult than it first appears. First, the link can be text-based, image-based, or both. For example, there is a lexicon of phrases used to indicate print-friendly pages, such as "print", "print article", "print-friendly version", etc. In addition, some links use printer-resembling image icons with or without a print phrase present. To complicate the matter further, not all the links contain a valid URL, but instead the pages are dynamically generated either by the client Javascript or by the server, so no URL is available for extraction. We estimate that there are more than 90% of the Web article pages have print-links, of which about 35% of them have valid print-friendly URLs, which is a good percentage. Our solution to the print-link extraction problem takes on two stages: (1) the detection of the print-link, (2) the retrieval of the print-friendly page URL from the link attributes, including the test for its validity. Experimental results based on roughly 2000 web article pages suggest our solution is capable of achieving over 99% precision and 97% recall performance measures.

Read the paper · More papers on PaperTik