ClinicalLayoutLM: A Pre-trained Multi-modal Model for Understanding Scanned Document in Electronic Health Records

Qiang Wei, Xu Zuo, Omer Anjum, Yan Hu, Ryan Denlinger, Elmer Victor Bernstam, Martin J. Citardi, Hua Xu · 2022 IEEE International Conference on Big Data (Big Data) · 2022

Scanned documents (e.g., faxes) are still widely used in clinical practice and are prevalent in Electronic Health Records (EHR). Unlocking information in scanned documents in EHRs is critical for clinical operation and research. However, it is challenging as it requires converting images to texts before applying information extraction technologies. Here we propose a multi-modal approach (ClinicalLayoutLM) that jointly models text extracted from Optical Character Recognition (OCR) and layout/image information to classify scanned clinical documents into different categories (e.g., lab reports and CT scans). Using a clinical corpus of 348, 311 scanned documents, we continually pretrained ClinicalLayoutLM based on LayoutLMv3, a multi-modal model from the open domain. For the task to classify the scanned clinical documents into 16 categories, ClinicalLayoutLM achieved an F1 score of 0.9051, which outperformed the baseline model (0.8840) that was based on text from OCR only. ClinicalLayoutLM is the first of its kind of multi-modal models for clinical documents and we believe it could benefit other clinical natural language processing (NLP) tasks such as layout analysis, information extraction and so on. The code is available at https://github.com/UTHealth-CCB/ClinicalLayoutLM and the pre-trained model is available upon request.

Read the paper · More papers on PaperTik