A Resource-Light Approach to Morpho-Syntactic Tagging Anna Feldman* and Jirka Hana‡ (*Montclair State University, ‡Charles University) Amsterdam: Rodopi (Language and computers: Studies in practical linguistics, volume 70), 2010, xiv+185 pp; hardbound, ISBN 978-90-420-2768-8, €40.00

Christian Monson · Computational Linguistics · 2011

Anna Feldman and Jirka Hana had a problem.Wanting to extract Russian verb frames, they lacked a tool for the necessary first step: morphological analysis of Russian words, disambiguated for context.To avoid the significant overhead of building a contextualized morphological analyzer from scratch, Feldman and Hana wondered if an analyzer that was already available for Czech would perform adequately on Russian.This book is the culmination of five years' research on projecting to a target language a contextualized morphological analyzer that was built for a separate source language, when both source and target belong to the same language family (Slavic, Romance, etc.).The authors succeed at building competitive morphological analysis systems for the target languages they consider (Russian, Catalan, and Portuguese), while expending a minimum of effort to construct specialized resources for these targets.At the culmination of their book, in Chapter 7, Feldman and Hana report a 6% absolute improvement, 79.7% vs. 73.5% labeling accuracy, when using a Czech morphological analyzer projected to Russian as opposed to training a statistical analyzer directly on a small sample (1,758 words) of hand-annotated Russian.Unfortunately missing is a formal demonstration that hand-labeling 1,758 words with morphological analyses requires an equivalent human effort to projecting an analyzer from one language to another.The authors' final morphological projection incorporates a variety of improvements that require human intervention: from a handbuilt morphological guesser on the target language side, to hand-defined rules that identify cognates between source and target languages and that render the syntactic structure of the source language more similar to the target's structure.Nowhere do the authors report the person-hours required to build each of these components and the reader is left to trust that constructing the projected systems takes as little time as is implied.A word of warning to those with a linguistics background: The authors prefer the language of natural language processing (NLP) to standard linguistic terminology.As a prime example, the title of this book includes the phrase morpho-syntactic tagging, a term from NLP. Part-of-speech tagging, in languages with little inflectional morphology, such as English, involves assigning to each word one part-of-speech tag from a small set of 50 or so possible tags.For the more inflected Slavic and Romance languages considered in this book, the tag sets include as many as 4,000 tags, each marking a full suite of morphosyntactic features, such as tense, case, or number.Thus, in linguistic terms, morphosyntactic tagging is exactly contextually disambiguated morphological analysis.Feldman and Hana's book substantiates a number of useful insights beyond the core observation that projecting morphology to a target language may be both adequate and cheaper than starting from scratch.The authors show, for example, that even in languages with significant inflectional morphology, such as Czech, the vast majority of

Read the paper · More papers on PaperTik