Determining the Validity of Clustering for Data Fusion
Kate Smith‐Miles, Sheldon Chuan, Peter van der Putten · 2002
For many direct marketing activities, organisations frequently find that customer databases do not contain enough information. Additional databases such as socio-economic databases constructed from census and survey data can be purchased to supplement customer databases. One of the difficulties in fusing separate databases however is that the information is based on two different samples and rarely can a unique individual be identified in both samples. Usually a common set of variables are used to determine the similarities between customers in the two samples, and various methods have been proposed for then predicting the missing information from one sample based on the information contained in the other sample. While some complicated methods have been proposed for data fusion, in this paper we investigate the validity of a simple clustering approach. Using a set of variables common to both samples, clusters are generated based on the k-means algorithm. The likely values of missing variables are then inferred based on the average values within the relevant cluster. An out-of-sample test set is used to demonstrate the accuracy of the fused variable predictions. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.