Adaptive Post Recognition
Philipp Berger, Patrick Hennig, Dominic Petrick, Marcel Pursche, Christoph Meinel · 2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2014) · 2014
Blogs, news portal and discussion forums are of high interest for today's social interaction research. But the automatic information extraction from the raw html page of those media channels is still a well-known problem. We introduce a novel approach to infer website templates based on the syndication format of blogs and news portals, called feeds. In comparison to related approaches that infer templates by clustering generic pages, we do not rely on a manual annotated training set. Instead, we use the feeds and their linked articles as training set to identify characteristic XPaths. Those paths identify the exact article content and article properties like title, author and publishing date. Further, we can use those paths to detect article pages that are no longer linked from feeds. We show the precision gain by comparing the article content extraction with an alternative approach e.g. boilerplate.