LANGUAGE IDENTIFICATION IN COMPLEX, UNORIENTED, AND DEGRADED DOCUMENT IMAGES
Dar-Shyang Lee, C.R. Nohl, Henry S. Baird · Series in machine perception and artificial intelligence · 1998
this report, combine decisions (or statistics) from all the text lines on the page to support a final decision at the page level. The order of execution is as follows. Within a page, we visit each text line (in arbitrary order), extracting features and computing statistics for each (details of these will be given below). Each text line is immediately classified as Asian or Latin (or rejected), in a way that requires no prior knowledge of the orientation. Feature statistics which support other decisions are accumulated and, at the end of the page, the remaining decisions are made in the following order. First, we distinguish between Asian and Latin scripts, determined by a majority vote of the text-line classifications already decided. If the page's script is Asian, we then decide the page's language using feature statistics accumulated from all text lines (again, at any orientation); then, using knowledge of the Asian-script page's language, we detect the page's orientation. If the page's script is Latin, we detect the page's orientation first and then decide the page's language. An overview of this process is shown in Figure 1. Details of the methods used for each of these decisions, along with test results, are discussed in the sections that follow. But first, we briefly summarize the training and test data used. 3 Training and Testing Data