User-based document clustering by redescribing subject descriptions with a genetic algorithm
Michael D. Gordon · Journal of the American Society for Information Science · 1991
Information retrieval systems have used clustering of documents and queries to improve both retrieval efficiency and retrieval effectiveness. Normally, clustering involves grouping together static descriptions of documents by their similarity to each other, though user-based clustering suggests that usage patterns concerning co-relevance can form a basis for clustering. This article reports that clusters of co-relevant documents obtain increasingly similar descriptions when a genetic algorithm is used to adapt subject descriptions so that documents become more effective in matching relevant queries and failing to match nonrelevant queries. As a result of the increased similarity, clustering algorithms can more accurately group documents into useful clusters. The findings of this work were reached through simulation experiments. © 1991 John Wiley & Sons, Inc.