Statistical recognition of content terms in general text

Martin C. Dillon, Peggy Federhart · Journal of the American Society for Information Science · 1984

Abstract This article discusses ways to improve the quality of retrieval systems that depend on the use of truncated words or quasi‐word stems as an indexing vocabulary. The problems addressed are the generalizability and stability of discriminant function analysis for selecting good topical terms from terms of relatively high frequency in a database drawn from abstracts of Harris Survey press releases. Results confirm that topical terms can be identified by their statistical properties. Consistently high recall of topical terms under a variety of different conditions implies persistent underlying properties strong enough to resist changes in test environment.

Read the paper · More papers on PaperTik