Document image segmentation using averaging filtering and mathematical morphology
Marina V. Polyakova, Alesya V. Ishchenko, Natallia Huliaieva · 2018 14th International Conference on Advanced Trends in Radioelecrtronics, Telecommunications and Computer Engineering (TCSET) · 2018
Scanned document image segmentation into text and non-text regions is an important preprocessing step before optical character recognition and document compression. Document image segmentation approaches perform pixel-based, block-based, or connected component processing. Block-based segmentation approaches depend on the accuracy of block partitioning. Pixel-based image segmentation is time-consuming. The execution time and performance of connected component based approaches depend on a number of connected components which is defined by half-toning. This article is devoted to modification of Bloomberg's approach with hole-filling by using linear average filtering instead threshold reduction. As a result a percentage of non-text pixels classified as text is decreased.