Syntax-Based Collocation Extraction Violeta Seretan (University of Geneva) Berlin: Springer (Text, speech and language technology series, volume 44), 2011, xi+217 pp; hardbound, ISBN 978-94-007-0133-5, $139.00

Pavel Pecina · Computational Linguistics · 2011

Collocation is a common language phenomenon which has attracted the interest of researchers in many subfields of both theoretical and computational linguistics.Although there is no commonly accepted and precise definition of this phenomenon, collocations are generally understood as complex lexical items, often characterized as unpredictable, idiosyncratic, holistic, mutually selective, and so forth.Together with other types of multiword expressions (or phraseological units, such as compound nouns, phrasal verbs, idioms, etc.), collocations form a borderline phenomenon positioned between lexis and grammar: On one hand, they are unpredictable and must be learned in the same way as single words are (as whole units); on the other hand, they often also have internal syntactic structure and their components must then adhere to grammatical rules.Collocations play an important role in applications involving text production (e.g., machine translation and language generation), text analysis (e.g., parsing and word sense disambiguation), and also in other related tasks (such as information extraction, text classification, etc.).The book Syntax-Based Collocation Extraction by Violeta Seretan is based on her doctoral dissertation defended in 2008 at the Department of Linguistics, University of Geneva, under the supervision of Eric Wehrli, and refers to a number of their previous publications.The main text is divided into six chapters (amounting to 128 pages) and six appendices (70 pages).The first chapter can be regarded as a motivation for the whole work.It introduces the notion of collocation, explains its relevance (and importance) for natural language processing, specifies the aims of the work, and most importantly, it presents arguments for syntax-based collocation extraction as a more appropriate alternative to the traditional syntax-free n-gram and window-based techniques.

Read the paper · More papers on PaperTik