Information retrieval from heterogeneous data sets using moderated IDF-cosine similarity in vector space model
Bhagyashree Pathak, Niranjan Lal · 2017
In a digital world information retrieval used for several application such as searching information in a document, searching for documents themselves, searching for metadata that depict data and for databases such as text, HTML, XML, image, audio etc. many universities, schools and companies use this IR systems to provide access to books, journals and other documents required in companies. Search engines are the most observable IR application. Our purpose is to compose system which search in our collection of datasets and retrieve the document which matches the user query. These collections of documents are text files, PDF's, HTML, XML files which we call heterogeneous data sets. For getting relevant documents in final outcomes, first we need to index the documents then rank them according to the score of cosine similarity values for each document which matches to the user query that we are entering. We are using cosine similarity technique in vector space model mainly for our analysis purpose and comparing with moderated IDF-Cosine similarity for better searching results using Rstudio. Our approach gives better results on real world corpus as comparing with enhanced IDF-Cosine similarity approach in vector space model.