A novel Arabic lemmatization algorithm

Eiman Tamah Al-Shammari, Jessica Lin · 2008

Tokenization is a fundamental step in processing textual data preceding the tasks of information retrieval, text mining, and natural language processing. Tokenization is a language-dependent approach, including normalization, stop words removal, lemmatization and stemming.

Read the paper · More papers on PaperTik