Information Extraction: Algorithms and Prospects in a Retrieval Context Marie-Francine Moens (Katholieke Universiteit Leuven) Springer (Information retrieval series, edited by W. Bruce Croft), 2006, xiii+246 pp; ISBN 978-1-4020-4987-3, $119.00

Diana Maynard · Computational Linguistics · 2008

Published as part of the Information Retrieval Series, this book aims to present both a historical overview of information extraction (IE) and a description of current approaches and applications.Essentially, it introduces the topic of information extraction to an information retrieval (IR) audience, aiming as much at students as established researchers, and focusing primarily on the algorithms used.Although it claims to give equal importance to early technologies developed in the field and to the "most advanced and recent technologies," for the latter it concentrates mainly on machinelearning approaches rather than knowledge-engineering (rule-based) techniques.The book is well-structured and progresses from an introductory chapter that gives a brief explanation of information extraction and how it fits into the IR paradigm, through a historical overview of the field in Chapter 2, and on to the meat of the book, which describes different approaches to IE in a set of four chapters leading from symbolic techniques and pattern recognition through to supervised and then unsupervised classification techniques.Chapter 7 discusses information retrieval models and makes some suggestions as to why and how information extraction should be incorporated into these models.Chapter 8 then discusses evaluation of IE technologies, focusing on the most commonly used measures from the Message Understanding Conference (MUC) and Automatic Content Extraction (ACE) competitions, and suggesting other ways in which evaluation might be measured.Chapter 9 describes some case studies in which information extraction is commonly used.Finally, the author takes a look at the future of information extraction within an information retrieval context, discussing some of the findings and challenges for current research.The aim of the book is quite ambitious in attempting to cater to the rather different needs of both students and established researchers.On the one hand, the chapters on the history of information extraction would be interesting for students; on the other hand, this could be of lesser interest to researchers concerned about the techniques and applications.The first half of the book is quite readable; the second half, however, targets those with quite some knowledge already of statistical processing, language modeling, and so on.Although the book is designed more for reading right through than dipping into individual chapters, the middle sections are certainly not easy reading, except perhaps for those already familiar with the topic.It is slightly disappointing that the book is angled almost exclusively towards machine-learning approaches to IE, and consequently omits any mention of some of the leading tools in the field such as GATE (Cunningham et al. 2002) and KIM (Popov et al. 2004), which are based on knowledge-engineering approaches.Although the book does not concern itself with examining individual IE tools, one might expect some reference to such tools, at least in the historical chapters if nowhere else.It is also unfortunate

Read the paper · More papers on PaperTik