Improving Search and Retrieval Performance through Shortening Documents, Detecting Garbage, and Throwing Out Jargon

Scott Kulp · 2007

This thesis describes the development of a new search and retrieval system used to index and process queries for several different data sets of documents. This thesis also describes my work with the TREC Legal data set, in particular, the new algorithms I designed to improve recall and precision rates in the legal domain. I have applied novel normalization techniques that are designed to slightly favor longer documents instead of assuming that all documents should have equal weight. I have created a set of rules used to detect errors in text caused by optical character recognition software. I have also developed a new method for reformulating query text when background information is provided with an information request. Using the queries and judgment data released by TREC in 2006, query pruning with cube root normalization results in average recall improvement of 67.3% and average precision improvement of 29.6%, when compared to cosine normalization without query reformulation. When using the 2007 queries, extremely high performance improvements are found using different kinds of power normalization, with recall and precision improving by over 400 % in the top rankings. Chapter 1

Read the paper · More papers on PaperTik