Determining the Number of Data Clusters in Any Dataset in the Presence of Considerable Noise

S. Easwaran · ASME Press eBooks · 2007

In this paper, a technique to determine the number of data clusters in a dataset in the presence of considerable extraneous noise data is proposed. The technique is based on intelligently combining the seed-growing algorithm for number-of-clusters determination with the weighted voting-scheme for analyzing data clusters. The generally satisfactory performance of this technique is seen through a variety of artificially created datasets intentionally corrupted with considerable noise data. Some issues with this technique and situations where this algorithm may not perform too well are also discussed.

Read the paper · More papers on PaperTik