Lexicalisation Patterns in Danish and Spanish
Casper A. G Woldersgaard · 2017
In this dissertation, I investigate the theoretical framework by Leonard Talmy (2000b) on lexicalisation patterns in Motion events. I examine his characterisations of Co-event languages (e.g., Danish) and Path-event languages (e.g., Spanish), and I relate his work to a Danish language setting. Furthermore, my objective is to determine whether the predictions set forth by Talmy apply to Danish and Spanish from an empirical perspective, i.e., in a Danish monolingual reference corpus, Korpus-DK, and a Spanish monolingual reference corpus, CORPES. I present different methods for testing Talmy’s theory when applied to large quantities of linguistic data. The methods I propose stem from corpus linguistics and natural language processing. I analyse the Danish corpus in order to present and exemplify a set of methods from corpus linguistics, e.g., collostructional analysis and the use of dispersion measures. However, the suggested methods do present shortcomings. For example, collostructional analysis requires a very clean input in the sense that false matches need to be manually »weeded out« before running the analysis. In corpora, there are vast amounts of linguistic data available. As a consequence, to identify Motion events and discard false positives is an extremely time-consuming process. I suggest that a context-free grammar is a way to facilitate the retrieval and analysis of linguistic data that contain Motion events. Thus, I implement a context-free grammar for Spanish. More specifically, I develop and formalise a language model that, when handed over to a computer program (a parser), will automatically retrieve and analyse (parse) sentences that contain Motion events. The advantage with this type of language model is that you need to account for the whole range of possibilities that a language has in its inventory to express Motion. Otherwise, the retrieved data will be noisy, and the proposed analyses will contain errors. At the same time, this is a disadvantage because the process of implementing a language model becomes time-consuming. Even though the set of methods are described separately in order to compartmentalise the overall picture into text, the methods show their true strength when they are used jointly. Because they are language independent, the methods can be applied to a wide range of languages. Hence, they will enable us to arrive at a more precise and empirically grounded characterisation of how languages express Mottion and, ultimately, augment and refine the findings in Talmy’s framework.