An optimal feature selection method for approximately duplicate records detecting
Hua Quan-ping, Ming Xiang, Fangyi Sun · 2010
During duplicate records detection and recognition in large number of data sets, detection accuracy is low and cost of detecting is high because that source of data are complicated and there are too many feature attributes. To solve these questions, we proposed an optimal feature selection method based on fuzzy clustering in groups. First, it deals with attributes of records in groups so as to reduce dimensions of attributes recorded effectively and obtain representative records in groups. Then it detects approximately duplicate records in groups by a computing method which compares with similarity. With theory analysis and experiments, it shows that identification accuracy and detection efficiency of this method are higher and it can solve recognition problem of approximately duplicate records in large number of data sets better.