An approach to address identification from degraded address data
Viresh Seth · afips · 1899
Today's Optical Character Recognition (OCR) technology does not read characters with 100 percent accuracy. Thus, the data string read by OCR may have one or more of the following deficiencies.• Unrecognizable characters• Incorrectly read characters• Added characters• Missing characters• Subclass charactersThis degradation then poses special problems in performing contextual analysis on address data strings comprising the identification of relevant address elements and the comparison of these address elements with entries in an address directory. Having some knowledge of the nature of the data, context analysis first attempts to correct one or more of the above mentioned deficiencies. Next a search for known keywords such as street designators (Avenue, etc.) is performed on the data. Finding keywords helps identification of the position of the other address elements and the determination of the type of address element. A search of the relevant files in the address directory then yields a definite address identification. However, due to the degraded nature of the data, keywords either escape detection or are erroneously found which creates confusion in contextual analysis.The comparison of the address data with entries in the directory is performed in the hardware primarily because of real time considerations. The algorithm used is based on a weighting technique which compensates for the deficiencies in the data. Generally, the algorithm is successful in reducing the comparison of ten or less potential candidates from the directory. Then, contextual analysis in the software attempts to isolate the unique candidate yielding the correct sort information for this mail piece.