A HYBRID METHOD OF LINGUISTIC APPROACH AND STATISTICAL METHOD FOR NESTED NOUN COMPOUND EXTRACTION

Hamed Al-Balushi · 2014

Arabic noun compound extraction has become a challenging issue in the field of NLP. Several approaches have been proposed in terms of extracting Arabic noun compounds. Some of them have used linguisticbased approach, other have used statistical methods and the rest have used a hybrid between them. This research proposed a hybrid method of linguistic-based approach and statistical method in order to solve the extraction of nested Arabic noun compound. The dataset has been collected from online Arabic newspaper archive from Aljazeara.net and Almotamar.net. Several pre-processing steps have been carried out on the data including transformation, normalization, stemming and POS tagging. After that, the n-gram model has generated bi-gram, tri-gram, 4-gram, and 5-gram candidates of noun compound. Then three association measures which are NC-value, PMI and LLR have been used in order to rank the candidates. The evaluation has been performed using the n-best method with a human annotation (manual selection by expertise). NC-value has outperformed PMI and LLR in terms of extracting nested noun compounds.

Read the paper · More papers on PaperTik