Evaluating linguistic knowledge in neural networks
Geoffrey Iain Bacon · eScholarship (California Digital Library) · 2020
Where does knowledge of language come from? How, for example, do speakers learn the meanings of words or the restrictions on their co-occurrences? This age-old question has age- old answers, from the necessity of direct sensory experience of the world to the existence of an innate language faculty. Recently, neural networks trained on distributional data have proven enormously successful in applied natural language processing tasks, suggesting that they acquire substantial knowledge of language. This dissertation examines what neural networks learn about language. Specifically, I present four studies that characterize the phonological, morphosyntactic and semantic knowledge of neural networks across more than 80 languages. The first of these focuses on phonological features and I show that distributional data of modest size is sufficient to induce human-like phoneme representations using standard neural architectures. The second uses agreement relations as a means of assessing sensitivity to structure dependence in a state-of-the-art model. Using a new cross-linguistic dataset of four types of agreement relations, I demonstrate that the model does capture syntax-sensitive agreement patterns well in general, but I also highlight the specific linguistic contexts in which its performance degrades. The third study looks at the lexical semantics of visual concepts in two domains, comparing neural models to both sighted and blind speakers’ representations. These analyses show that some human-like knowledge is captured, but that the more nuanced structures of the domains are not. Taken together, these first three studies argue that neural networks trained on distributional data are largely accurate yet imperfect models of language. The final study of this dissertation suggests a way forward. In this study, I show that the semantic typology of tense systems is well explained by a domain-general pressure for communicative efficiency and suggest that this same principle is an appropriate inductive bias for neural networks, which may lead to developing more human-like computational models of language.