Natural language processing for learner corpus research

Kristopher Kyle · International Journal of Learner Corpus Research · 2020

The term natural language processing (NLP) refers to the use of computer programs to automatically analyze human language.NLP processes range from the (relatively) simple task of splitting character sequences into words and sentences to much more sophisticated (and challenging) tasks such as converting speech sounds into text and annotating texts for syntactic, semantic, and pragmatic features (among others, see Jurafsky & Manning, 2008 for a survey of common NLP processes; and Meurers & Dickinson, 2017 for specific applications to L2 research).NLP tools of varying complexity have played an important role in the development of corpus linguistics in general and learner corpus research (LCR) in particular.Although relatively simple NLP tools such as concordancers (e.g., AntConc; Anthony, 2019; Wordsmith Tools; Scott, 2020) and related programs (AntWordProfiler; Anthony, 2014; VocabProfile; Cobb, 2018; Range; Heatley & Nation, 1994) have been used extensively in the field of LCR, advances in machine learning 1 have made much more complex analyses possible.Part of speech (POS) taggers, such as TreeTagger (Schmid, 1994), CLAWS (Garside, Leech, & McEnery, 1997), and the Stanford POS Tagger (Toutanova, Klein, Manning, & Singer, 2003), for example, automatically annotate texts with POS tags, allowing for more finegrained analyses than is possible with unannotated texts (e.g., Bestgen & Granger, 2014;Biber, Gray, & Staples, 2014;Granger & Bestgen, 2017).Syntactic parsers such as MaltParser (Nivre, Hall, & Nilsson, 2006), the Stanford Parser (Chen & Manning, 2014; Klein & Manning, 2003), and spaCy (Explosion AI, 2018) automatically annotate texts for syntactic constituency or dependency relationships.Syntactic parsers allow researchers to automatically investigate even more complex linguistic features such as dependency bigrams (e.g., Kyle & Eguchi, in press;

Read the paper · More papers on PaperTik