The effect of disfluencies and learner errors on the parsing of spoken learner language

Andrew Caines, Paula J. Buttery · 2014

NLP tools are typically trained on written data from native speakers. However, research into language acquisition and tools for language teaching & proficiency assessment would benefit from accurate processing of spoken data from second language learners. In this paper we discuss manual annotation schemes for various features of spoken language; we also evaluate the automatic tagging of one particular feature (filled pauses) ‐ finding a success rate of 81%; and we evaluate the effect of using our manual annotations to ‘clean up’ the transcriptions for sentence parsing, resulting in a 25% improvement in parse success rate by completely cleaning the texts of disfluencies and errors. We discuss the need to adapt existing NLP technology to non-canonical domains such as spoken learner language, while emphasising the worth of continued integration of manual and automatic annotation.

Read the paper · More papers on PaperTik