Constraint-based incremental clustering algorithm with mixed attributes
Renxia Wan · Jisuanji gongcheng yu sheji · 2010
To solve the constraint of the memory capacity during clustering the large-scale dataset, a fast clustering algorithm based on the constraint of the number of clusters is put forward. The original dataset is read only once and the radius threshold changes dynamically. At the same time an inter-cluster dissimilarity measure taking into account the frequency information of the categorical attribute values is introduced, which can be used for the mixed dataset. The time complexity and space complexity are nearly linear with the size of dataset and the number of attributes. The experimental results on the KDDCUP99 dataset show that the proposed algorithm is feasible and effective, which can be used for the large-scale dataset.