Improvements in predicting children's overall reading ability by modeling variability in evaluators' subjective judgments

Matthew P. Black, Shrikanth Shri Narayanan · 2012

Automatic literacy assessment is one promising application of speech and language processing research. In our previous work, we showed we could accurately predict children's overall ability to read a list of English words aloud, an integral component of early literacy assessment. In this paper, we improve upon our results by exploiting the fact that evaluators' level of agreement significantly varies, depending on the child being judged. This source of evaluator variability is directly modeled using generalized least squares linear regression. In this framework, the children for which the evaluators were more confident in rating are weighted higher. Performance in predicting the mean evaluator's scores increases from a Pearson's correlation coefficient of 0.946 to 0.952, a relative improvement of 0.63%. This is a significantly higher correlation than the mean inter-evaluator agreement of 0.899 (p <; 0.05). Critically, the mean and maximum absolute errors are significantly reduced.

Read the paper · More papers on PaperTik