Tony McEnery and Andrew Hardie. Corpus Linguistics: Method, Theory and Practice.
Adam Kilgarriff · International Journal of Lexicography · 2012
This book gives a beautifully clear account of where corpus linguistics is today. It addresses those issues that lurk behind any corpus research: sampling, corpus types, corpus-based vs. corpus-driven, copyright (for texts) and privacy (for interviews, conversations etc), annotation, statistics, and whether we select data for our analyses, or whether we are duty bound to look at everything that our corpus routines throw up. It is a textbook. Each chapter starts with an introduction and ends with a summary plus sections of ‘Further Reading’, ‘Practical Activities’ and ‘Questions for Discussion’. It is explicitly not an introductory textbook: as the authors note, several exist already1 and this does not aim to compete with them. The target audience is those who have already worked with corpora and have found themselves encountering the challenges and theoretical perplexities that the book addresses. The Introduction gives a sketch of each of the issues, plus the basic contrasts of spoken vs. written, and monolingual vs. multilingual. The second chapter covers some history, corpus analysis tools, and statistics. The history section covers the Chomsky objections to corpora, but does not linger on them, which I appreciated: they are well-worn territory, and not especially current.