Automatic lemmatization of Persian words*
Tayebeh Mosavi Miangah · Journal of Quantitative Linguistics · 2006
This study presents a rather novel method for suffix and prefix stripping of Persian words. The method presented is a language independent one and mostly relies on a specially arranged corpus composed of a list of roots, word-forms, prefixes, and suffixes which has been manually compiled. Applying an algorithm, the morphemes of each word including root, prefix, and suffixes are separated from each other specifying their place and order in the given word. The degree of accuracy and precision of this program is dependent on the representative characteristics and richness of the corpus. The present method for which Persian words have been tested represents an accuracy of 87.24% in Persian. Although the main application of this algorithm is in the field of information retrieval, it can be used in a machine translation system from Persian into any other language. In this case a stem dictionary for morphological analysis should be used.