Stemming algorithm for different tenses to improve Persian dictionary

Arash Ghazvini, Mohd Juzaiddin Ab Aziz · 2012

Persian language is an Indo-European language that is known for its complexity due to the morphology structure. In this paper, we report on Persian stemmer and the impact on improvement of Persian dictionary. Persian language consists of a variety of tenses, while the focus is on past subjunctive, past perfect, continuous past, present perfect and past simple. In Persian language, it is important to get rid of affixes from the verbs to obtain the stem. Therefore, finite state machine has been chosen to develop a Persian stemmer. According to the findings and testing results, Persian stemming algorithm based dictionary is fully accurate for the regular verbs in mentioned tenses.

Read the paper · More papers on PaperTik