Automatic Learning of Stemming Rules for the Indonesian Language
Lily Suryana Indradjaja, Stéphane Bressan · Institutional Repositories DataBase (IRDB) · 2003
We present a method for the automatic learning of stemming rules for the Indonesian language.The learning process uses an unlabelled corpus.In the first phase the candidate (word, stem) pairs are automatically extracted from a set of online documents.This phase uses a dictionary but is nevertheless not trivial because of morphing.In the second phase the rules are induced from the thus obtained list of pairs of words with their respective stems.We evaluate the effectiveness of our method for different sizes of the training set with different settings on the thresholds of the support and confidence of each rule.We discuss how these variables affect the quantity and quality of the rules produced.