Web page classification based on context to the content extraction of articles

Ankit Dilip Patel, Vimal N. Pandya · 2017

Now-a-days, World Wide Web is growing as a huge repository with high source of voluminous and heterogeneous information which continues to expand in size and complexity. But still Web pages are observed unstructured, semi-structured or asynchronous, which contains excess irrelevant information like advertisement, sponsored links, headers, footers etc. Hence guiding on to fetch particular documents over the Internet is well supported by search engines, on giving appropriate keywords in form of queries, or either by catalogues generation, which organize documents into hierarchical file structure. But maintenance of such catalogues manually is more difficult, due to the huge data residing on the Web; hence it becomes necessary build some techniques for auto-categorization of documents. Auto-categorizations scales the retrieval of more relevant information by crawling through web page contents and comparing with the keywords which exact matches the behavioral of any particular category. The prime objective of this research is to identify the category in which the article falls, which enrich the reader's knowledge with direct access the content or locate to the relevant titles. The paper describes a model to perform categorization on article related to child development and parenting's contexts, which starts with catalog generation through identification & analysis carried with the consideration of category keyword along with relevant information been extracted from various sources over web to achieve the classification with expected accuracy.

Read the paper · More papers on PaperTik