On the Reliability and Validity of Human and Lsa-Based Evaluations of Complex Student-Authored Texts
Eva Seifried, Wolfgang Lenhard, Herbert Baier, Birgit Spinath · Journal of Educational Computing Research · 2012
This study investigates the potential of a software tool based on Latent Semantic Analysis (LSA; Landauer, McNamara, Dennis, & Kintsch, 2007) to automatically evaluate complex German texts. A sample of N = 94 German university students provided written answers to questions that involved a high amount of analytical reasoning and evaluation. LSA-based evaluations were compared to evaluations of six human graders. Results showed that LSA-based evaluations agreed as well with human graders as human graders agreed with each other. Agreement of human graders and LSA-based evaluations did not differ depending on the standard of comparison used by LSA, that is, whether scoring was based on one single text (“gold standard”) or on a sample of previously scored assignments (“nearest neighbors”). Moreover, LSA-based evaluations of students' assignments predicted students' results in a final exam. These results show that LSA can support university teachers in giving feedback and grading complex German texts.