An Approach to Content Extraction from Scientific Articles using Case-Based Reasoning
Rajendra Prasath, Pınar Öztürk · Research in Computing Science · 2016
In this paper, we present an efficient approach for content extraction of scientific papers from web pages.The approach uses an artificial intelligence method, Case-Based Reasoning(CBR), that relies on the idea that similar problems have similar solutions and hence reuses past experiences to solve new problems or tasks.The key task of content extraction is the classification of HTML tag sequences where the sequences representing navigation links, advertisements and, other non-informative content are not of interest when the goal is to extract scientific contributions.Our method learns from each experience with the tag sequence classification episode and stores these in the case base.When a new tag sequence needs to be classified, the system checks its case base to see whether a similar tag was experienced before in order to reuse it for content extraction.If the tag sequence is completely new, then it uses the proposed algorithm that relies on two assumptions related to the distribution of various tag sequences occurring in the page, and the similarity of the tag sequences with respect to their structure in terms of levels of the tags.Experimental results show that the proposed approach efficiently extracts content information from scientific articles.