Can we utilize Large Language Models (LLMs) to generate useful linguistic corpora? A case study of the word frequency effect in young German readers

Job J. Schepens, Hanna Woloszyn, Nicole Marx, Benjamin Gagl · 2023

Can a measure of word frequency (i.e., how often a word occurs) based on text generated by a large language model (LLM) model the well-studied word frequency effect in reading? To study this, we created a large corpus of text generated by an LLM (GPT 3.5) to measure word frequency. We prompted it to write stories in German directed to children using the titles of 500 popular books, so that the corpus would follow the structure of an existing corpus based on books written for children (childLex). The generated text had a lower lexical richness than the text in childLex. We then used reading performance data from German children (i.e., reaction times in a lexical decision task) to investigate the effect of LLM word frequency in visual word recognition. We found that LLM word frequency explained more variance in reaction times when compared to the word frequency measure based on the original children's books (Experiment 1) and also when compared to a word frequency measure based on adult-directed LLM text (Experiment 2). We conclude that LLM corpora are less lexically rich while approximating children's lexical processing well. We discuss the potential of this approach, considering the inevitable risks of using LLMs that lack openness, are highly complex, and need a large amount of resources.

Read the paper · More papers on PaperTik