Border Detection for Camera-Captured Document Images Using Transformers
Mahita Kandala, Kaushik M, Vignesh G S, Peeta Basa Pati · 2024
In the era of digitization, accurately delineating the boundaries of camera-captured documents stands as a crucial yet challenging task, impacting various processes such as Optical Character Recognition (OCR) and document analysis. This paper presents a novel approach to document boundary detection using SegFormer, an innovative model leveraging Transformers' capabilities for contextual understanding. The research aims to address the limitations of traditional methods by harnessing Segformer's ability to capture intricate spatial relationships within images. In addition, a comparative analysis with two state-of-the-art models, UNet and DeepLabV3+, is conducted to evaluate the effectiveness of SegFormer in document boundary detection tasks. The experimentation involves training and testing the models on a dataset of camera-captured document images, meticulously annotated and segmented to ensure robust performance across diverse layouts. Through this research, valuable insights into the application of Transformers in document boundary detection are contributed, highlighting the importance of context-aware models in segmentation tasks.