Text region extraction method for historical Tibetan document based on border detection

Yiqun Wang, Weilan Wang, Zhenqi Cai · 2022

The text region extraction is the key step for optical character recognition (OCR) operation. According to the layout characteristics of historical Tibetan document, this paper proposes a text region extraction method based on border detection to extract text region from documents. Firstly, the character height and stroke width are estimated and the border region position is detected. Then, the position of decorative lines surrounding the body text region is detected by heuristic search. Finally, the mask image of the document is generated according to the position relationship between the text regions and border, then different text regions are extracted. Experiments on dataset of historical Tibetan document show that this method can effectively overcome the problems of page tilt, border and decorative lines fracture, and demonstrate the effectiveness of the proposed method.

Read the paper · More papers on PaperTik