How well can a corpus-derived co-occurrence network simulate human associative behavior?

Gemma Bel-Enguix, Reinhard Rapp, Michael Zock · 2014

Free word associations are the words people spontaneously come up with in re-sponse to a stimulus word. Such informa-tion has been collected from test persons and stored in databases. A well known example is the Edinburgh Associative Thesaurus (EAT). We will show in this paper that this kind of knowledge can be acquired automatically from corpora, en-abling the computer to produce similar associative responses as people do. While in the past test sets typically consisted of approximately 100 words, we will use here a large part of the EAT which, in to-tal, comprises 8400 words. Apart from extending the test set, we consider differ-ent properties of words: saliency, fre-quency and part-of-speech. For each fea-ture categorize our test set, and we com-pare the simulation results to those based on the EAT. It turns out that there are surprising similarities which supports our claim that a corpus-derived co-occur-rence network can simulate human asso-ciative behavior, i.e. an important part of language acquisition and verbal behavior. 1

Read the paper · More papers on PaperTik