Classifying easy-to-read texts without parsing
Johan Falkenjack, Arne Jönsson · 2014
Document classification using automated linguistic analysis and machine learning (ML) has been shown to be a viable road forward for readability assessment. The best models can be trained to decide if a text is easy to read or not with very high accuracy, e.g. a model using 117 parame-ters from shallow, lexical, morphological and syntactic analyses achieves 98,9 % ac-curacy. In this paper we compare models created by parameter optimization over subsets of that total model to find out to which extent different high-performing models tend to consist of the same parameters and if it is possible to find models that only use fea-tures not requiring parsing. We used a ge-netic algorithm to systematically optimize parameter sets of fixed sizes using accu-racy of a Support Vector Machine classi-fier as fitness function. Our results show that it is possible to find models almost as good as the currently best models while omitting parsing based features. 1