Image sampling rate and image pattern recognition

Geetha Srikantan · 1994

The effect of dimensionality of data representation on the performance of a pattern classification system has been studied by several researchers. Theoretical results relate performance and measurement dimensionality for finite and infinite training set sizes. Previous work has focussed on general measurement representations. This dissertation examines the relation between the data representation determined by the sampling rate and the recognition accuracy in pattern classification systems. Increasing the sampling rate of digitization of data is modeled here as a refinement of the representation, without requiring reconstructibility of data digitized at a lower sampling rate. Previous results bounding the probability of error of a pattern classification system are examined under the refinement constraint. As the sampling rate is increased, performance improvements are expected; this is reflected in the improved bounds on the probability of error under the refinement constraint. While theoretical bounds on error probability are of great interest, the assumptions of data set size under which they hold are difficult to satisfy in practice. In practical problems such as optical character recognition (or OCR), this problem takes on a particular significance. The OCR system performance time can be improved considerably by operating on coarse data representations (i.e. low sampling rate of digitization); also for certain problems, high resolution data may not be available. An estimate of the error probability under these conditions is needed. A computational model is proposed to examine the effect of sampling rate on optical character recognition. A novel application of multirate filter theory for resampling images by fractional factors is presented. Several feature extraction and classification methods have been used in the experiments. English and Japanese handwritten and machine-printed characters of varying quality have been utilized in the experiments. The empirical results show that increasing the sampling rate indeed, improves recognition accuracy across all data sets. Conversely, the results also provide an estimate of the intrinsic error in recognition for coarse data representation.

Read the paper · More papers on PaperTik