Feature Analysis on English word difficulty by Gaussian Mixture Model

Hua Yang, Suyong Eum · 2018

Machine Learning has significantly improved Natural Language Process (NLP) recently. In this paper, we firstly adopt the NLP approaches to extract features which represent the difficulty levels of English words. Then, Principal Component Analysis (PCA) is applied to reduce the feature dimension for visualization of data points, which intuitively justifies the selection of the features. More elaborated analysis is carried out using Gaussian Mixture Model (GMM) which clusters the data points in the reduced dimensional space. The analysis verifies that the proposed features appropriately contribute to the prediction of English word difficulty level. Finally, we demonstrate that 73.5% of classification accuracy can be achieved with Support Vector Machine (SVM) with the features we proposed.

Read the paper · More papers on PaperTik