Kazakh Noun Phrase Extraction Based on N-gram and Rules

Gulila Altenbek, Ruina Sun · 2010

The aim of the work is to extract Kazakh phrase and basic noun phrase from corpus. For the phrase extraction, N-gram model methods were used, specifically bigram and trigram methods were applied. For basic noun phrase extraction, rule-based methods were used. We started from the grammar structure of basic noun phrase structure model, established a set of rules using the part-of-speech tag and the additional component information of Kazakh basic noun phrase, and extracted the basic noun phrase by rule matching. We have realized the extraction of phrase and basic noun phrase based on corpus of 31 days' Xinjiang Daily. Experimental results showed that the two methods are feasible, and the extraction accuracies are 50.8% and 79.1% respectively.

Read the paper · More papers on PaperTik