Using self-organizing recognition as a mechanism for rejecting segmentation errors
Charles L. Wilson · 1992
We have developed a self-organized neural network based method that concurrently de- tects segmentation errors and performs character recognition.This method utilizes a two-pass classification scheme.A page of macliine printed text is segmented, and a pre-trained self- organizing classifier is used to recognize the images produced by the segmenter.Images that are recognized with a sufficiently high confidence axe used to retrain the classifier, adapting the neural network to the current font type being segmented.All the segmented images are then reclassified by the adapted network.The assigned classes of those images which are confidently recognized are accepted, whereas the images which are not confidently recognized are rejected.The pages of text used to develop this method were randomly generated so that no context can be used to correct segmentation or recognition errors.In one experiment, the first classification pass rejected 6.6% of the total number of images segmented.Of these rejected images, only 3.6% were truly segmentation errors.The other 96.4% were correctly segmented.A traditional single-pass classification scheme such as this results in the rejection of an unnecessarily high number of correctly segmented characters, reducing effective system throughput.By retraining the network on only those images accepted by the first pass, the second classification pass rejected only 3.5% of the segmented images, of which 6.7% were segmentation errors.The second pass achieved an accuracy of 99.3% on those images accepted.This clearly demonstrates the network's ability to adapt on the second pass, increasing system throughput by 3.1%.In all the cases studied, greater than 99% of all segmentation errors were detected without any human intervention.1