DATA MINING WITH NATURAL LANGUAGE PROCESSING AND CORPUS LINGUISTICS

Alison L. Bailey, Anne Blackstock‐Bernstein, Ève Ryan, Despina Pitsoulakis · 2016

This chapter brings together the fields of corpus linguistics, natural language processing (NLP), and computing to describe how language samples of school-age students were used to create a digital data system of dynamic language learning progressions (DLLPs). It focuses on the small number of studies examining language corpora applied to school-age children's formal education and provides some of the motivations and design decisions underlying the DLLP project. The chapter highlights the convergence of corpus linguistics, NLP, and data mining and their key characteristics and how these components combine to be applied to the analysis, assessment, and instruction of language. It addresses the instabilities and dysfluencies of young children's speech and the developmental idiosyncrasies of their oral and written productions. The chapter emphasizes the dearth of research on and practices with language corpora in the education of young language learners.

Read the paper · More papers on PaperTik