Learning with Mislabeled Training Samples Using
Amita Pathak-Pal, Sankar Kumar Pal · 1987
For the problem of parameter learning in pattern recognition, the convergence of stochastic approximation-based learning algorithms have been investigated for the situation in which mislabeled training samples are present. In the cases considered, it is found that estimates converge to nontrue values in the presence of labeling errors. The general m-class A r -feature pattern recognition problem is considered. A possible solution to the problem is also discussed. Some simulation results are provided to support the conclusions drawn. I. INTRODUCTION The learning of unknown parameters of classifiers is an indis- pensible part of pattern recognition problems. If a sufficiently large set of correctly labeled training samples is available, then reasonably good estimates of the parameters can generally be obtained. In many real-life situations, however, it is either diffi cult or expensive to obtain labels, so that mislabeling of training samples can become one of the specters with which a pattern recognition scientist has to contend. It is, therefore, useful to know how this problem can affect the learning procedure. A reasonable amount of work has been done for the two-class classification problem. The effects of random training errors on Fisher's discriminant function have been studied by Lachenbruch (1), (2). McLachlan (3), Michalek and Tripathi (4), O'Neill (5), Krishnan (6), and Katre and Krishnan (7). They concluded that the effect is to underestimate distance, overestimate error rate, introduce bias into estimates of the discriminant function, make the maximum likelihood estimates of the discriminant function converge to nontrue values, and change the asymptotic relative efficiency (ARE) relative to a completely correctly classified sample of the same size. In the context of recursive learning of parameters, the useful ness of stochastic approximation procedures cannot be overem phasized (8). 1 Briefly, a stochastic approximation procedure for recursively estimating a parameter Θ by θ (at the n th stage) with the help of an unbiased statistic T is x n ~~ ~~ X n ~~ *n * 1 ) »