RobustLayoutLM: Leveraging Optimized Layout with Additional Modalities for Improved Document Understanding
Bowen Wang, Xiao Wei · 2024
Pre-training methods have become the mainstream in document understanding, involving self-supervised learning on large-scale unlabeled document to learn rich feature representations, followed by supervised fine-tuning on a smaller labeled document dataset. Recently emerged pre-training models for document understanding all use layout information as important features, however, their performance significantly suffers when the layout information is incorrect. In this paper, we introduce RobustLayoutLM, which utilizes the optimized layout information generated by our proposed Uniform XY Cut algorithm. And Visual-Graph Tagger is introduced to integrate image and graph information after the transformer feature fusion, minimizing the impact of erroneous layout information. Experimental results show that our RobustLayoutLM achieves competitive or better results on multiple standard VrDU benchmarks and outperforms previous methods in the face of incorrect layout information.