Evaluating Modern Information Extraction Techniques for Complex Document Structures

Harichandana Gonuguntla, Kurra Meghana, Thode Sai Prajwal, Koundinya N V S S, Kolli Nethre Sai, Chirag Jain · 2024

The exponential increase in digital data has necessitated the development of advanced models for Key Information Extraction (KIE) from text-heavy documents. Traditional Optical Character Recognition (OCR) systems have evolved to incorporate multimodal vision approaches, enhancing their capacity to extract not just text but also tables, key-value pairs, figures, and other significant document components. This paper provides a comprehensive review of the state-of-the-art models including Visual Information Extraction, Semantic Entity Recognition (SER), Relation Extraction (RE), and Document Layout Analysis. Despite considerable advancements, these methods still encounter various challenges that undermine their effectiveness and efficiency. This paper presents a comparative study of five different information extraction approaches applied to a carefully curated dataset of text-heavy documents, including various types such as forms, invoices, and long-form legal and medical reports. Our findings demonstrate that this integrated OCR and multimodal vision approach achieves superior accuracy compared to other approaches. We present a critical analysis of the challenges faced in this domain and propose suggestions for future research to refine information extraction capabilities.

Read the paper · More papers on PaperTik