Att-CapsViT: Attention-based CapsNet and Vision Transformer Hybride Network for Handwriting Chinese Characters Recognition
Xiangling Zheng, Nanxiang Zhou, Zibin Geng · 2024
The research on text recognition has long been under the close scrutiny of scholars, especially the recognition of complex Chinese characters, which is highly favored by the academic community. In recent years, handwriting recognition has achieved remarkable results within the traditional framework of deep learning. However, this approach still falls short in capturing the subtle features and spatial hierarchical relationships within Chinese characters. Furthermore, the complexity of Chinese handwritten characters and the diversity of writing styles pose higher challenges to models. To address this, we propose an attention-based CapsN et and vision transformer hybride network for handwriting Chinese characters recognition (Att-CapsViT), which combines the attention mechanism with capsule networks and visual transformer. The Att-CapsViT model integrates the attention mechanism, combining the local detail feature capturing capability of capsule networks (CapsNets) with the global contextual awareness of visual transformer (ViT) to achieve multi-level feature extraction from local to global. This experiment was validated using two open-source datasets, successfully confirming the robust performance of the Att-CapsViT model in the recognition of Chinese handwritten characters, providing a new perspective for the field of complex text recognition.