Taking a new look at the latent semantic analysis approach to information retrieval

Elizabeth R. Jessup, James Martin · 2001

Latent Semantic Analysis (LSA) [4] is a mathematical approach to the discovery ofsimilarity relationships among documents, fragments of documents, and the wordsthat occur within collections of documents. Although LSA was originally appliedin the context of information retrieval [4], it has since been successfully applied toa wide variety of text-based tasks [16].LSA is a variant of the vector space model for information retrieval that usesa reduced-rank approximation to the term-document matrix. In the informationretrieval domain, rank reduction is applied in an e ort to remove the oise thatobscures the semantic content of the data [4]. In this context, two claims aretypically made for LSA: that it provides a substantial improvement in retrievalperformanceoverthestandardvectorspacemodelandthatthisimprovementresultsfrom LSA’s ability to solve what is known as the synonymy problem.Despite the many successful applications of LSA, there are a large numberof unanswered questions that bear on where, and in what manner, LSA should beapplied. The purpose of this paper is to begin to investigate these issues in LSA’soriginal context of information retrieval and to pose new directions for future work.Among the more critical questions that we address in this paper are the following: Does LSA reliably improve retrieval performance as compared to the vectorspace model? Does LSA improve retrieval performance by addressing the synonymy prob-lem? How can the optimal rank be chosen? How can relevant and irrelevant documents be distinguished? And are there alternative matrix techniques that can be used to discoverreduced representations?This paper is organized as follows. In Sections 0.2{0.3, we review the details of thevector space model and LSA. In Section 0.4, we outline our empirical methods. InSection 0.5, we compare the retrieval performances of LSA and the full-rank vectorspace model. In Section 0.6, we evaluate how the performance of LSA depends onits ability to handle synonyms. In Sections 0.7{0.8, we consider the choice of rankand how best to identify relevant documents. In Section 0.9, we examine the useother orthogonal decompositions for rank reduction. Finally, in Section 0.10, wesummarize our results.

Read the paper · More papers on PaperTik