A Hybrid Method of Syntactic Feature and Latent Semantic Analysis for Automatic Arabic Essay Scoring

R. Mezher, Nazlia Omar · Journal of Applied Sciences · 2016

Background: The process of automated essays assessments is a challenging task due to the need of comprehensive evaluation in order to validate the answers accurately.The challenge increases when dealing with Arabic language where, morphology, semantic and syntactic are complex.Methodology: There are few research efforts have been proposed for Automatic Essays Scoring (AES) in Arabic.However, such efforts have concentrated on the semantic perspective by proposing Latent Semantic Analysis (LSA).The LSA is based on word-document co-occurrence, also called a ʻBag-of-wordsʼ approach.It is therefore blind to the syntactic information.This puts limitations on LSAʼs ability to capture the meaning of a sentence which depends upon both syntax and semantic.Therefore, using syntactical features may improve the process of evaluation.Hence, this study proposed a hybrid method of syntactic features and LSA for automatic essay scoring.Several pre-processing tasks have been performed in order to normalize the words with an appropriate format for processing.Then, the similarity matrix of LSA will be constructed using Term Frequency-Inverse Document Frequency (TF-IDF).After that, the cosine similarity will be carried out to identify the similarity among words.Results: Finally, part of speech (POS) tagging is applied in order to identify the syntactic feature of words within the similarity matrix.The dataset contains 61 questions related to environmental science with 10 answers for each question in, which the total number of answers is 610.Conclusion: The experimental results have shown that the syntactic feature improves the accuracy of AES for Arabic language.

Read the paper · More papers on PaperTik