Extracting web content for personalized presentation

Rodrigo Chamun, Daniele Pinheiro, Diego Jornada, João Batista S. de Oliveira, Isabel Harb Manssour · 2014

Printing web pages is usually a thankless task as the result is often a document with many badly-used pages and poor layout. Besides the actual content, superfluous web elements like menus and links are often present and in a printed version they are commonly perceived as an annoyance. Therefore, a solution for obtaining cleaner versions for printing is to detect parts of the page that the reader wants to consume, eliminating unnecessary elements and filtering the "true" content of the web page. In addition, the same solution may be used online to present cleaner versions of web pages, discarding any elements that the user wishes to avoid.

Read the paper · More papers on PaperTik