Parameter optimization for BIRCH pre-clustering algorithm
László Kovács, László Bednarik · 2011
The pre-clustering is an efficient data reduction method in the case of large data sets. In the pre-clustering process, an important aspect is to provide a good intra-cluster similarity. Most of the traditional methods do not consider this aspect and they generate weak clusters. The paper presents some algorithms to optimize the key parameters (like branching factor, quality threshold and selection of the separator line) of the BIRCH pre-clustering method.