Machine-Aided Linguistic Discovery: An Introduction and Some Examples Vladimir Pericliev (Bulgarian Academy of Sciences) London: Equinox, 2010, ix+330 pp; hardbound, ISBN 978-1-84553-660-2, $90.00, £60.00

Eric Smith · Computational Linguistics · 2010

The subtitle of Vladimir Pericliev's book, An Introduction and Some Examples, is a succinct and accurate description of its contents.Pericliev argues briefly for the usefulness of computer-aided techniques in linguistic discovery, contrasting it with the intuitionist approach which has characterized linguistic discovery throughout much of its history.The bulk of the book is devoted to examples of software-aided linguistic discovery drawn from his own work.Chapter 1 starts by sketching out the current state of discovery techniques in linguistic theory, categorizing scientific discovery into three main approaches: the intuitionist approach, the chance approach, and the problem-solving approach.Discoveries by intuition and by chance remain the purview of humans, but clearly the problemsolving approach can benefit from the application of computational techniques.Chapter 2 presents the KINSHIP program, which performs "parsimonious discrimination" in order to determine the minimal set of features which are necessary to discriminate all of a language's kinship terms.The program is used to discover feature geometries, superior to existing human-discovered ones, which describe the kinship terminology of languages like English and Bulgarian.Chapter 3 extends the ideas used in KINSHIP to a program called MPD (maximal parsimonious discrimination), which is then applied to a variety of other tasks, some of which are unconnected to linguistics.Of these applications, the most interesting is the use of MPD to determine the segment profiles which uniquely identify languages in the UPSID-451 database (consisting of segment inventories from 451 languages, selected to provide broad coverage of the world's language families) (Maddieson and Precoda 1991).Although Pericliev discusses his results at considerable length, it is not clear what the theoretical usefulness of these profiles might be.What does it really tell us about French to know that it is the only language in the database to contain the phoneme [ œ]?Of more practical interest was Pericliev's discussion of the process of converting the UPSID data into a featural representation to make it amenable to processing, describing how to represent underspecified segments and how to deal with transcription variations.This sort of necessary preprocessing constitutes an important and underemphasized part of the process of machine-aided linguistic discovery.The study of UPSID does produce some interesting, though not unexpected, results.For instance, when a profile contains more than one unique segment, the majority of these segments share a common feature, and 85.8% of the unique segments have some sort of secondary articulation.Chapters 4 and 5 present two more pieces of software developed by Pericliev: UNIV and AUTO.The UNIV software is inspired by Greenberg's universals (Greenberg 1966),

Read the paper · More papers on PaperTik