Complex words identification using word-level features for SemEval-2020 Task 1

Jenny Ortiz Zambrano, Arturo Montejo‐Ráez · 2021

This article describes a system to predict the complexity of words for the Lexical Complexity Prediction (LCP) shared task hosted at Se-mEval 2021 (Task 1) with a new annotated English dataset with a Likert scale.Located in the Lexical Semantics track, the task consisted of predicting the complexity value of the words in context.A machine learning approach was carried out based on the frequency of the words and several characteristics added at word level.Over these features, a supervised random forest regression algorithm was trained.Several runs were performed with different values to observe the performance of the algorithm.For the evaluation, our best results reported a M.A.E of 0.07347, M.S.E. of 0.00938, and R.M.S.E. of 0.096871.Our experiments showed that, with a greater number of characteristics, the precision of the classification increases.

Read the paper · More papers on PaperTik