Topic-Specific Scoring of Documents with Discrete PCA ?

Wray Buntine, Kimmo Valtonen · 2005

The random surfer model for scoring of documents, for in- stance using PageRank, works when good link structure exists for a col- lection. Here, we develop a topic-specific version using a topic struc- ture developed automatically via discrete PCA methods. To evaluate the resultant method, scores are developed on the Wikipedia, the pub- lic domain encyclopedia on the web, because it has a good internal link structure, and results can be readily interpreted from the page titles. More sophisticated language models are starting to be used in information retrieval (8) and real successes are being achieved in their use (4). A document modelling approach based on discrete versions of PCA (7,1,2) has been applied to the language modelling task in information retrieval (2,3). Here we apply the same discrete PCA method to topic specific versions of page rank (6,10). Our intent is that this can be used as a secondary score to add topical scoring to retrieval in conjunction with a separate key-word based score such as TFIDF.

Read the paper · More papers on PaperTik