Automatic evaluation of spoken summaries: the case of language assessment

Anastassia Loukina, Klaus Zechner, Lei Chen · 2014

This paper investigates whether ROUGE, a popular metric for the evaluation of automated written summaries, can be applied to the assessment of spoken summaries produced by non-native speakers of English.We demonstrate that ROUGE, with its emphasis on the recall of information, is particularly suited to the assessment of the summarization quality of non-native speakers' responses.A standard baseline implementation of ROUGE-1 computed over the output of the automated speech recognizer has a Spearman correlation of ρ = 0.55 with experts' scores of speakers' proficiency (ρ = 0.51 for a content-vector baseline).Further increases in agreement with experts' scores can be achieved by using types instead of tokens for the computation of word frequencies for both candidate and reference summaries, as well as by using multiple reference summaries instead of a single one.These modifications increase the correlation with experts' scores to a Spearman correlation of ρ = 0.65.Furthermore, we found that the choice of reference summaries does not have any impact on performance, and that the adjusted metric is also robust to errors introduced by automated speech recognition (ρ = 0.67 for human transcriptions vs. ρ = 0.65 for speech recognition output).

Read the paper · More papers on PaperTik