An Effective Pre-Processing Algorithm for Information Retrieval Systems

Vikram Singh, Balwinder Saini · International Journal of Database Management Systems · 2014

The Internet is probably the most successful distributed computing system ever.However, our capabilities for data querying and manipulation on the internet are primordial at best.The user expectations are enhancing over the period of time along with increased amount of operational data past few decades.The data-user expects more deep, exact, and detailed results.Result retrieval for the user query is always relative o the pattern of data storage and index.In Information retrieval systems, tokenization is an integrals part whose prime objective is to identifying the token and their count.In this paper, we have proposed an effective tokenization approach which is based on training vector and result shows that efficiency/ effectiveness of proposed algorithm.Tokenization on documents helps to satisfy user's information need more precisely and reduced search sharply, is believed to be a part of information retrieval.Pre-processing of input document is an integral part of Tokenization, which involves preprocessing of documents and generates its respective tokens which is the basis of these tokens probabilistic IR generate its scoring and gives reduced search space.The comparative analysis is based on the two parameters; Number of Token generated, Pre-processing time.

Read the paper · More papers on PaperTik