Comprehensive Evaluation of Transformer Models for Complex Information Detection in Unstructured Documents
Divya Singh, Atsumi Terayama · 2024
Information detection from unstructured data, such as address detection, is crucial in many applications. Traditional approaches like regular expressions (RegEx) are computationally efficient but struggle to adapt to the variability in data formats, leading to low accuracy. In contrast, machine learning-based approaches, particularly those involving transformer-based models, have shown great promise in natural language processing due to their adaptability and high accuracy. This paper provides a comprehensive evaluation of several prominent transformer-based models Bidirectional Encoder Representations from Transformers (BERT), DistilBERT, and RoBERTa—in the context of address detection. BERT, known for its deep contextual understanding, offers high accuracy but comes with a higher computational cost. DistilBERT, a distilled version of BERT, provides a faster and more resource-efficient alternative but may sacrifice some accuracy and RoBERTa is optimized for performance. Experiments, conducted on a dataset combining the National Address Database with Wikipedia texts, revealed that RoBERTa outperformed the others, achieving 99% accuracy with a 0.0193-s inference time. In contrast, using RegEx achieved less than 10% accuracy. This evaluation highlights the effectiveness of transformer-based models, particularly RoBERTa, in detecting complex information such as addresses, while also considering resource consumption. The findings suggest that RoBERTa, strikes the best balance between accuracy and computational cost among the models tested, making it a strong candidate for broader applications in unstructured document analysis.