Developing a persian chunker using a hybrid approach
Soheila Kiani, Tara Akhavan, Mehrnoush Shamsfard · 2009
Text segmentation is the process of recognizing boundaries of text constituents, such as sentences, phrases and words. This paper focuses on phrase segmentation also known as chunking. This task has different problems in various natural languages depending on linguistic features and prescribed form of writing. In this paper, we will discuss the problems and solutions especially for the Persian language and present our system for Persian phrase segmentation. Our system exploits a hybrid method for automatic chunking of Persian texts. The method at first exploits a rule-based approach to create a tagged corpus for training a neural network and then uses a multilayer perceptron neural network and Fuzzy C-Means Clustering to chunk new sentences. Experimental results show the average precision of %85.7 for the chunking result.