Persian Text Normalization using Classification Tree and Support Vector Machine
Mohammad Hossein Moattar, Mohammad Mehdi Homayounpour, Davoud Zabihzadeh · 2006
Text normalization is one of the most important tasks in text processing and text to speech conversion. In this paper, we propose a machine learning method to determine the type of Farsi language non-standard words (NSWs) by only using the structure of these words. Two methods including support vector machines (SVM) and classification and regression trees (CART) were used and evaluated on different training and test sets for NSW type classification in Farsi. The experimental results show that, NSW type classification in Farsi can be efficiently done by using only the structural form of Farsi non-standard words. In addition, the results is compared with a previous work done on normalization using multi-layer perceptron (MLP) neural network and shows that SVM outperforms MLP in both number of efforts and total performance