Quantitative Linguistic Computing with Perl. Fengxiang Fan, Yaochen Deng.
Lei Lei · Literary and Linguistic Computing · 2012
The book under review introduces fundamentals of Perl programming for quantitative linguistic studies. Though the authors explicitly claim that the target audience of the book is ‘broad-spectrum’ (p. 3), the book may be of most help to those linguists who have no programming background, or those who are familiar with other programming languages and would like a quick reference guide to the Perl language. Perl is a scripting language which is free, cross-platform, and extremely powerful in natural language processing. Thanks to these features, it is widely used in quantitative, computational, and corpus linguistic studies. However, as the authors point out, Perl code may not be easy to read and understand. Another weakness of Perl I would like to address here is that the learning curve of the language is rather steep, especially to those who have no programming experience. For such readers, this book will serve well as a ‘stepping stone’ (p. 2) to the skills of Perl programming, as ‘no previous computer programming experience is required for this book’ (p. 3). The authors have done a good job to achieve this goal. First, the book provides step-by-step instructions on the programming elements and skills and gradually unfolds to its reader the mystery of Perl programming. It starts from the most elementary basics, such as installation of Perl, Perl code execution, major data structures and their processing, and progresses to more advanced topics, such as regular expressions, subroutines and modules, file and folder management etc. Furthermore, it addresses problems commonly encountered in quantitative, computational, and corpus linguistic research, such as concordancing, word list construction, and lexical bundles, or Ngrams extraction. Additionally, the book also provides solutions to some topics of interest in quantitative linguistics such as lemmatization, word entropy, word frequency spectrum, syllabic word length computation, Chinese tokenization, and hapax legomena. Last, most chapters are accompanied with a section of exercises, which provide the readers with opportunities to practice the theoretical concepts discussed in the chapters concerned.