Ranking of text documents using TF-IDF weighting and association rules mining
Siham Jabri, Azzeddine Dahbi, Taoufiq Gadi, Abdelhak Bassir · 2018 4th International Conference on Optimization and Applications (ICOA) · 2018
The amount of information has increased exponentially over the past few years. Most of the text documents are unstructured and not organized and hence user facing more difficult to find his information needs. It creates a challenge for information retrieval for ranking documents according to the user relevance. One of more useful and efficient retrieval model is the vector space model. It uses term weighting tf-idf for assigning a score to a query/document pair, in order to rank the documents. This method is based on the bag-of-words model and does not capture position of terms in document and semantics. Keeping this approach, here we propose a new ranking measure that combines the vector space measure denoted tf-idf and new factor deducted from association rules technique based on dominance relations. This approach will help the user to get the most relevant documents at the beginning. The experiments on TREC collection show that the proposed method is giving better results compared to the baselines and the semantic ranking ESA.