Vision Mamba-Based Approach for Incomplete Boundary Document Image Rectification

Weihao Zhang, Xin Tao Xia, Maopeng Li, Yun‐Bo Zhao · 2025

Capturing document images using handheld mobile devices often results in geometric deformations, which adversely affect the accuracy of Optical Character Recognition (OCR) and document understanding. However, existing transformer-based methods face significant computational costs when processing document images on resource-constrained devices. This study proposes an enhanced Vision Mamba architecture to learn the structural information of document images, thereby rectifying deformed images while reducing computational resource consumption. Additionally, owing to the relative positioning between the document and the imaging device, captured images may exhibit incomplete boundaries. Conventional learning-based methods are primarily designed for images with complete boundaries, which can diminish correction effectiveness. To address this issue, mask consistency loss and preprocessing techniques are introduced to improve the rectification of document images with incomplete boundaries. Experimental results demonstrate the effectiveness and superiority of this method, highlighting its significant value for intelligent document processing.

Read the paper · More papers on PaperTik