Receipt Recognition Technology Driven by Multimodal Alignment and Lightweight Sequence Modeling
Jin-Ming Yu, Huijun Ma, Jianlei Kong · Electronics · 2025
With the rapid advancement of global digital transformation, enterprises and financial institutions face increasing challenges in managing and processing receipt-like financial documents. Traditional manual document processing methods can no longer meet the demands of modern office operations and business expansion. To address these issues, automated document recognition systems based on computer vision and deep learning technologies have emerged. This paper proposes a receipt recognition technology based on multimodal alignment and lightweight sequence modeling, integrating the CLIP (Contrastive Language-Image Pretraining) and Bidirectional Gated Recurrent Unit (BiGRU) framework. The framework aims to achieve synergistic optimization of image and text information through semantic correction. By leveraging dynamic threshold classification, geometric regression loss, and multimodal feature alignment, the framework significantly improves text detection and recognition accuracy in complex layouts and low-quality images. Experimental results show that the model achieves a detection F1 score of 93.1% and a Character Error Rate (CER) of 5.1% on the CORD dataset. Through a three-stage compression strategy of quantization, pruning, and distillation, the model size is reduced to 18 MB, achieving real-time inference speeds of 25 FPS on the Jetson AGX Orin edge device, with power consumption stabilized below 12 W. This framework provides an efficient, accurate, and edge-computing-friendly solution for automated receipt processing. Practical implications include its potential to enhance the efficiency of financial audits, improve tax compliance, and streamline the operational management of financial institutions, making it a valuable tool for real-world applications in receipt automation.