Comparison of Stemming and N-gram Matching for Term Conflation in Arabic Text
Hani Abu‐Salem · International Journal of Computer Processing Of Languages · 2004
A stem in Arabic is a root verb form combined with derivational morphemes but with affixes removed. One part of this paper is to repeat the Mixed Stemming experiment, that chooses the use of a word, a stem, or a root for a query term based upon which form has the highest average inverse document frequency value, using more documents and queries. A consistency result has been obtained. The root with weighting method was the superior. The mixed stemming improved binary weighting search results in all cases but did not increase performance over weighted stems or roots. N-gram matching is widely used for term conflation when searching the World Wide Web for something without knowing whether it is stored as singular or plural, as a compound or in the form of several words, in old or new spelling or possibly in wrong orthography. Part two of this paper reports a comparison of Stemming and N-gram matching. The Stem-NGram-Stem Search outperforms (at recall 0.4 and above) the Digrams method when both the queries and documents words are not stemmed. It also suggests that the (Stem-NGram-Stem) Search approach outperforms (at recall 0.4 and above) the Trigrams method (Stem-NGram-Word) when the queries words are stemmed and documents words are not stemmed. The results also suggest that the Trigrams method (Stem-NGram-Word) outperforms the Digrams method when both the queries and documents words are not stemmed. All of the proposed N-gram methods outperform the Word, Stem, and Root index methods for Binary Weighting scheme.