Intelligent clustering with instance-level constraints

Kiri L. Wagstaff, Claire Cardie · 2002

One goal of research in artificial intelligence is to automate tasks that currently require human expertise; this automation is important because it saves time and brings problems that were previously too large to be solved into the feasible domain. Data analysis, or the ability to identify meaningful patterns and trends in large volumes of data, is an important task that falls into this category. Clustering algorithms are a particularly useful group of data analysis tools. These methods are used, for example, to analyze satellite images of the Earth to identify and categorize different land and foliage types or to analyze telescopic observations to determine what distinct types of astronomical bodies exist and to categorize each observation. However, most existing clustering methods apply general similarity techniques rather than making use of problem-specific information. This dissertation first presents a novel method for converting existing clustering algorithms into constrained clustering algorithms. The resulting methods are able to accept domain-specific information in the form of constraints on the output clusters. At the most general level, each constraint is an instance-level statement about a pair of items in the data set that indicates a preference for being placed into the same cluster, or, alternatively, into different clusters. The constrained clustering algorithms developed and presented in this dissertation enforce each constraint according to the strength of that preference. The second major contribution of this dissertation is the application of constrained clustering algorithms to diverse, significant, challenging real-world problems. We observe that the additional domain knowledge, when combined with the algorithms' ability to enforce that knowledge, produces improvements on a variety of tasks. The problem domains include automated map refinement, natural language processing, and automated data analysis of Hubble Space Telescope observations.

Read the paper · More papers on PaperTik