Semantic word sketches
Diana McCarthy, Adam Kilgarriff, Miloš Jakubíček, Siva Reddy · 2015
A central task of linguistic description is to identify the semantic and syntactic profiles of the words of a language: what arguments (if any) does a word (most often, a verb) take, what syntactic roles do they fill, and what kinds of arguments are they from a semantic point of view: what, in other terminologies, are their selectional restrictions or semantic preferences. Lexicographers have long done this ‘by hand’; since the advent of corpus methods in computational linguistics it has been an ambition of computational linguists to do it automatically, in a corpus-driven way, see for example (Briscoe et al 1991; Resnik 1993; McCarthy and Carroll 2003; Erk 2007). In this work we start from word sketches (Kilgarriff et al 2004), which are corpus-based accounts of a word’s grammatical and collocational behaviour. We combine the techniques we use to create these word sketches with a 315-million-word subset of the UKWaC corpus which has been automatically processed by SuperSense Tagger (SST) (Ciaramita and Altun 2006) to annotate all content words with not only their part-of-speech and lemma, but also their WordNet (Fellbaum 1998) lexicographer class. WordNet lexicographer classes are a set of 41 broad semantic classes that are used for organizing the lexicographers work. These semantic categories group together the WordNet senses (synsets) and have therefore been dubbed `supersenses' (Ciaramita and Johnson, 2003). There are 26 such supersenses for nouns and 15 for verbs.