SAID: A new stemmer algorithm to indexing unstructured Document
Kabil Boukhari, Mohamed Nazih Omri · 2015
In this work, we propose a new stemmer algorithm to indexing unstructured Document. It can detect the most relevant words in an unstructured document. This algorithm is based on two main modules: the first module ensures the processing of compound words and the second allows the detection of the endings of the words that have not been taken into consideration by the approaches presented in literature. The proposed algorithm allows the detection and removal of suffixes and enriches the basis of suffixes by eliminating the suffixes of compound words. We have experienced our algorithm on a standard basis of terms and the results show the remarkable effectiveness of our algorithm compared to others presented in related works.