Determining the Number of Data Clusters in Any Dataset in the Presence of Considerable Noise
S. Easwaran · ASME Press eBooks · 2007
In this paper, a technique to determine the number of data clusters in a dataset in the presence of considerable extraneous noise data is proposed. The technique is based on intelligently combining the seed-growing algorithm for number-of-clusters determination with the weighted voting-scheme for analyzing data clusters. The generally satisfactory performance of this technique is seen through a variety of artificially created datasets intentionally corrupted with considerable noise data. Some issues with this technique and situations where this algorithm may not perform too well are also discussed.