HUT-Net: A Hybrid U-Net Transformer Network with Attention Mechanisms for Historical Document Image Binarization

Sarra Hamza, Rafik Menassel, Chawki Djeddi · 2025

The degradation in quality and the complexity of document images significantly impact the preservation and analysis of historical manuscripts. In the interest of promoting accurate binarization of such document images, a novel deep learning architecture, HUT-Net, is presented: Hybrid U-Net Transformer with an Attention Mechanism. The proposed model tackles the local and global dependencies within complex document images effectively through a combination of strengths from transformer and convolutional architectures. To efficiently extract features, HUT-Net uses an EfficientNet-B0 architecture as the encoder backbone. A transformer block is applied at the bottleneck to aggregate global contextual information that is indispensable in discriminating text from background noises. In order to selectively refine characteristics and preserve fine textual nuances, the decoder incorporates skip connections and spatial attention methods. This hybrid strategy greatly improves the model's capacity to manage various degradations, including bleed-through, stains, and faint text, is much improved by this hybrid method. The model proved excellent performance on various measures, such as F1-Score, PSNR, and Distance Reciprocal Distortion (DRD), in evaluation tests performed over the widely known benchmark datasets, such as DIBCO 2017 and 2018. The comparative study highlights the strength of HUT-Net regarding currently proposed U-Net-based models, showing also robustness and adaptability for documents with diverse levels of degradation. The results demonstrate that HUT-Net is a strong and efficient method for binarization of historical documents, which will help the cultural heritage preservation field and increase the accessibility of digital archival materials.

Read the paper · More papers on PaperTik