Introduction to Information Retrieval Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze (Stanford University, Yahoo! Research, and University of Stuttgart) Cambridge: Cambridge University Press, 2008, xxi+482 pp; hardbound, ISBN 978-0-521-86571-5, $60.00

Olga Vechtomova · Computational Linguistics · 2009

Introduction to Information Retrieval by Manning, Raghavan, and Sch ütze is an up-todate, thorough, and systematic introduction to information retrieval (IR) from a computer science perspective.Written as a textbook, its main audience is graduate and senior undergraduate students taking IR courses.The book will also be valuable to researchers in other computer science fields, such as computational linguistics, as well as to professional practitioners wishing to delve into the IR field.The book is structured into 21 chapters, which gradually unfold the subject of information retrieval, starting with the fundamentals (such as Boolean retrieval, document indexing, vector-space model, and evaluation in IR) and moving on to more advanced topics (such as probabilistic models, XML retrieval, text classification, machine learning for IR, document clustering, and Web retrieval).Pedagogical features of the book include short exercises at the end of each section and brief overviews of related research literature at the end of each chapter.Some of the major strengths of the book are its accessibility, clarity, and good balance between theory and practice.There are many concrete examples throughout the book that facilitate understanding of complex topics.Although the book covers a broad selection of the major established and emerging topics in IR, it largely bypasses two important subjects, in my opinion: natural language processing techniques in IR and interactive information retrieval.Although the authors refer to some research done in these areas in various chapters, they do not give them the same thorough treatment given to other topics in the book.To compensate, in the preface the authors provide references to the detailed coverage of these and some other topics in other textbooks.It also might have been useful if the authors introduced some specialized IR tasks, such as opinion retrieval or enterprise search, which might benefit from more advanced NLP techniques.Chapter 1 gives a succinct and focused introduction to the main concepts in IR, such as term, index, document, query, recall, precision, and so on.It outlines the main principles of Boolean retrieval, briefly criticizes it, and compares it to ranked retrieval.The authors also present a good real-world example of a commercial Boolean retrieval system.Chapter 2 provides a detailed discussion of the initial stages of the document indexing process that include tokenization, stemming and lemmatization, stopwords removal, and approaches to dealing with phrases at the indexing stage, namely bigram indexing and the use of positional indexes.In this chapter the authors discuss some linguistic aspects of these processes.For example, when examining tokenization, they discuss various morphological and other aspects of languages that complicate this process (e.g., hyphenation in English, compound nouns in German, and word

Read the paper · More papers on PaperTik