Handbook of Natural Language Processing (second edition) Nitin Indurkhya and Fred J. Damerau (editors) (University of New South Wales; IBM Thomas J. Watson Research Center)Boca Raton, FL: CRC Press, 2010, xxxiii+678 pp; hardbound, ISBN 978-1-4200-8592-1, $99.95

Jochen L. Leidner · Computational Linguistics · 2011

The Handbook of Natural Language Processing is a revised edition of an earlier handbook (Dale, Moisl, and Somers 2000).This second edition was prepared by Nitin Indurkhya, a researcher at the University of New South Wales, and the late text processing pioneer Fred J. Damerau of the IBM T. J. Watson Research Center (d.27 January 2009), whose 1964 paper introduced a version of what is now known as the Damerau-Levenshtein distance, a metric of the similarity between two strings and a dynamic programming algorithm to compute it efficiently (Damerau 1964).Damerau also invented automatic hyphenation (Damerau 1970) and worked on early question-answering systems.Indurkhya, who is also affiliated with a consulting company, Data-Miner Pty Ltd., maintains a companion wiki for the book. 1 The book has three parts, totaling 26 chapters.The first part, Classical Approaches, essentially covers techniques that were known prior to the statistical revolution, that is, before natural language processing people in the mainstream embraced techniques that speech engineers were already using successfully for awhile.The second part, Empirical and Statistical Approaches, covers state-of-the-art data-driven models.2 Part three, Applications, shows some techniques closer to applications.If you are talking to a computational linguist, information extraction is seen as an application, but if you are talking to business people, they will see it as a general technology area, from which many application products and services can be built.The handbook "aims to cater to the needs of NLP practitioners and language-engineering professionals in academia as well as in industry. . . .The prototypical reader is interested in the practical aspects of building NLP systems and may also be interested in working with languages other than English" (p.xxii).Hence it would have been nice to introduce some descriptions of actual products that generate revenue (even if this meant that this particular part of the handbook would become outdated more quickly) in order to demonstrate how the NLP parts are embedded in non-NLP technology, and how these products are embedded in the businesses that use them.For example, the application chapter "Information Retrieval" does not describe how the topics in other parts were applied in Web search engines or enterprise search products, as one might have expected.Rather, it basically is another technical chapter-and its probabilistic IR material could just as well have been presented in part two (statistical techniques).

Read the paper · More papers on PaperTik