Implementation of PDF crawler using boolean inverted index and n-gram model
Snehal S. Kadwe, Shrikant B. Ardhapurkar · 2017
Today's world are mostly dependent on internet and electronic devices. Most of the users wish to store their information in PDF document, retrieval of such document are most formidable task. To overcome this problem, PDF crawler is implemented. PDF document can be retrieved using keyword and key-phrase present in it. The extraction of keyword is based on Boolean inverted index where as key-phrase is extracted using n-gram algorithm. The pre-processing of PDF document begins with assigning term frequency (TF) to each and every word available in it as well as each document is mapped with unique id called as (docID). After mapping the keyword with term-frequency it extract the keyword which has highest count and store into the database using inverted index with pair of docID and keyword. The key-phrase is extracted by using n-gram. Inverted index makes the pdf crawler faster by storing the documents at one place which contains the same keyword. It helps to reduce storage space as well as it optimized the time required to retrieve the document.