Evidence Extraction for Automated Medical Coding: Preliminary Evaluation

Xiaorui Jiang, Kulsoom Khan, Sumithra Thinakara Vasantha, Sajjad Ali Haider · 2024

Coding clinical texts in standard language such as ICD is an important but tedious and error-prone process.Automated medical coding algorithms suffer problems due to the combined the challenge of handling the significant length of clinical text, the complexity of the huge code hierarchy and the lack of interpretability to ensure user trust.Large language models (LLM) have also been proven struggling with this task in recent studies.Recent efforts have been made to annotate an evidence-supported medical coding dataset.The current study makes the first empirical investigation into how well (small) fine-tuned pretrained language models (PLM) and LLMs could identify the sentences containing medical evidence supporting the assigned codes.Hierarchical sequential sentence classification and GPT-3.5 in the zero-shot setting were tested for evidence sentence extraction.Extra evaluation was performed to investigate how evidence extraction impacts clinical coding and what implications it has towards the future generation algorithms for automated medical coding.

Read the paper · More papers on PaperTik