A comprehensive word difficulty index for L2 listening
Kourosh Meshgi, Maryam Sadat Mirzaei · ExLing Conferences · 2018
Word difficulty in listening tasks is considered challenging to assess because of its high subjectivity, high dimensionality, and low generalizability. We propose a word listening difficulty score formulated as a linear combination of several complementary features. A dataset of expert-annotated, partial, and synchronized captions for TED Talks was prepared for a target language proficiency level, in which only the difficult words are displayed. A linear Support Vector Machine (SVM) was trained on this dataset, and the learned parameters of the model were transferred to the proposed score. This data-driven score demonstrates higher accuracy on the annotated dataset and facilitates easy model and feature expansion.