Enriching the Extraction of Top-K Lists from the Web

Miss. Shreya Umesh Wadkar, Nilesh G. Pardeshi · International journal of advance research and innovative ideas in education · 2016

The web contains a tremendous amount of data and this data result in to big amount of information. This information on the web is of two type i) Structured data and ii) Unstructured data. In this, we concentrate on structure data. List data is most effective source of structure data for extracting the information from the web. This System deals with “Top-k Lists”, web pages that represent a list of k instances of a particular topic or concept. Examples are, “top 10 cricketers in the world”, “10 best dancer in the world” etc. Top-k lists are bigger, ranked and of good quality source of information. As a result, top-k lists are highly profitable .In this, we present an efficient method that obtain the target lists from web pages with high exactness. Compared to other structured data, top-k lists are clearer, easier to understand and more exciting for human use, and therefore are valuable source for knowledge mining and information finding. Extraction of such lists can help enrich existing knowledge bases about general concepts and significant as a pre-processing step to produce facts for a fact answering engine.

Read the paper · More papers on PaperTik