A neural network learning relative distances
Alfred Ultsch · 2000
Data mining and knowledge discovery aim at the detection of new knowledge in data sets produced by some data generating process. There are (at least) two important problems associated with such data sets: missing values and unknown distributions. This is in particular a problem for clustering algorithms, be it statistical or neuronal. For such purposes a distance metric comparing high dimensional vectors is necessary for all data points. Much handwork is necessary in today's data mining systems to find an appropriate metric. In this work a novel neural network, called ud-net is defined. Ud-nets are able to adapt to unknown distributions of data. The output of the networks may be interpreted as a distance metric. This metric is also defined for data with missing values in some components and the resulting value is comparable to complete data sets. Experiments with known distributions that were distorted by nonlinear transformations show that ud-net produce values on the transformed data that are very comparable to distance measures on the untransformed distributions. Ud-nets are also tested with a data set from a stock picking application. For this example the results of the networks are very similar to results obtained by the application of hand tuned nonlinear transformations to the data set.