2. A radically data-driven Construction Grammar: Experiments with Dutch causative constructions
Natalia Levshina, Kris Heylen · 2014
In this paper we propose a novel, radically data-driven approach to constructional semantics. It is based on Semantic Vector Space models, which are commonly used in computational linguistics to model the semantic relationships between words on the basis of their distribution in a large corpus. In a case study of the near-synonymous Dutch causative constructions with doen ‘do’ and laten ‘let’ we show this method in action by testing a variety of distributional models and clustering options applied to the constructional collexemes. The method opens new perspectives for generating hypotheses about constructional semantics, providing a quick estimation of large amounts of data. The paper also contributes to bridging the gap between the neostructuralist distributional approaches still predominant in computational linguistics, on the one hand, and the non-reductionist constructionist approaches to grammar, on the other hand. 1. The need for objective data-driven semantic classes Constructions are commonly defined as pairings of form and function (Goldberg 1995; Goldberg 2006). The meaning, understood here as the concept or conceptual structure associated with a construction, is a crucial aspect of the latter’s function. Although it is clear that the meaning of a construction cannot be reduced to the meaning of its components (e.g. Goldberg 1995), the semantic properties of its slot fillers can be used as a convenient heuristic to access the conventional uses of the construction in question. For instance, the central sense 1 This research project was partly funded by a grant from the Research Foundation of Flanders (FWO) (G.0330.08) awarded to Dirk Geeraerts and Dirk Speelman, the Quantitative Lexicology and Variational Linguistics Research Unit at the University of Leuven. Dr ft (c) Le vs hin a & He yle n