IMPROVEMENT OF ZONE CONTENT CLASSIFICATION BY USING BACKGROUND ANALYSIS
Yalin Wang, Ihsin T. Phillips · 2000
Abstract. This paper presents an improved zone content classification method. Motivated by our novel background-analysis-based table identification research, we added two new features to the feature vector from one previously published method [7]. The new features are the total area of large horizontal and large ver-tical blank blocks and the number of text glyphs in the zone. A binary decision tree is used to assign a zone class on the basis of its feature vector. The train-ing and testing data sets for the algorithm include images drawn from the UWCDROM-III document image database. The classifier is able to classify each given scientific and technical document zone into one of the nine classes, text classes (of font size pt and font size pt), math, table, halftone, map/drawing, ruling, logo, and others. The improved zone classification method raised the accuracy rate to fffi from \t \t flfi and reduced the median false alarm rate to ffi fi from fi. 1 Problem Statement Let! be a set of zone entities. Let " be a set of content labels, such as text, table, math, etc. The function #%$&!(') " associates each element of * with a label. The function +,$-!.'0 / specifies measurements made on each element of! , where / is the measurement space. The zone content classification problem can be formulated as follows: Given a zone set! and a content label set " , find a classification function #1$2!3'4 " , that has the maximum probability: