Discrepancy as a quality measure for avoiding classification bias
Kazunori Iwata, Naohiro Ishii · 2003
This paper discusses how to create initial data to achieve a good performance on active learning of multilayer perceptrons. The initial training data plays an important role for active learning performance, because any active learning algorithm generates additional training data based on existing data. In this paper on active learning of a multi-layer perceptron in the case of little initial data, we verify an effect of the bias of the initial data using discrepancy. Discrepancy is a measure of the uniformity of data distribution. We then discuss a method for generation of initial data using a low-discrepancy sequence. In our experimental results of the classification problem, we found that initial data with a low discrepancy avoids classification bias. Hence, discrepancy as a measure is a quality guide to avoid classification bias, and low-discrepancy sequences provide a good strategy to generate initial data on active learning of multi-layer perceptrons.