Using text-based indices, one can predict how lexically rich short texts in French, German, and Portuguese will be perceived
Rohit Bagthariya · IAAR Journal of Education - ISSN 2583-6846 Peer-Reviewed Journal · 2020
We investigated how accurately computerized calculations of the lexical features of the texts might predict readers' perceptions of the lexical richness of short texts. This was done by looking at the lexical characteristics of the texts. Over 150 indices were constructed from the writings of 3,060 children between the ages of 8 and 10 written in French, German, and Portuguese. They were assessed for their lexical richness by 3 and 18 raters who had not been educated. The writings ranged in length from 9 to 284 words. We observed that the precision with which the ratings of shorter texts could be expected was similar to that of longer texts and that the ratings could be predicted mainly based on these indices. However, models with fewer predictors based on a 6-dimensional framework of lexical richness perceptions or even just a straightforward predictor, Guiraud's index, performed slightly worse than models with more predictors. The most potent predictors for French and German were opaque models with scores of predictors.