A MultiModal Neural Network Architecture for Document Image Classification

S Krithika, A R Priyadharshini, G Bharathi Mohan, M. Gayathri · 2024

Document classification is a crucial domain in natural language processing that involves classifying text documents into predefined categories. However, in multiple cases, the documents are heterogeneous with different modalities such as images and flow diagrams, giving rise to the need for handling multiple modalities. Thus, multimodal document classification enables a more comprehensive understanding of document content by incorporating multiple modalities such as text and images. This enhances information extraction and enables more accurate categorization of documents, benefiting various applications such as content organization, search, recommendation systems, and information retrieval. This work proposes a deep learning approach that leverages both textual and image information for improved classification of multimodal documents. Our multimodal deep neural network architecture utilizes two different channels in processing the multimodal features: text and document images. The textual model leverages BERT to generate embeddings of the text, while the image feature extraction channel employs a hybrid architecture consisting of the pre-trained convolutional neural network model VGG-19 as the backbone followed by a vision transformer. The two different channels are incorporated into a unified framework that performs the final classification. The training and testing of the proposed framework were conducted on the benchmark TOBACCO-3482 dataset, demonstrating the effectiveness and potential of our proposed multimodal classification approach. The TOBACCO-3482 dataset is a widely used benchmark dataset for document image analysis and classification. The effectiveness in performing the classification of the proposed framework is evaluated on the grounds of performance metrics such as confusion matrix, accuracy, precision, recall, and F1-score.

Read the paper · More papers on PaperTik