Enhancing ECG Report Generation With Domain-Specific Tokenization for Improved Medical NLP Accuracy
Farzeen Ashfaq, Noor Zaman Jhanjhi, Navid Ali Khan, Danish Javed, Mehedi Masud, Mohammad Shorfuzzaman · IEEE Access · 2025
The automation of medical report generation has been an area of interest for researchers over the years, with significant advancements in computational techniques and natural language processing. Traditionally, the focus has remained on rule-based or template-driven approaches to streamline the documentation process. However, recent breakthroughs in generative AI models have substantially enhanced the capabilities of automated medical report generation. Hence, the field has gained significant attention in recent years due to its potential to support healthcare professionals and improve diagnostic workflows. One key challenge in these models is the reliance on pre-trained, general-purpose tokenizers, which often fail to capture the domain-specific vocabulary essential for accurate medical reporting. When using a general tokenizer, the generated reports may lack coherence, relevance, or even correct medical terminology, leading to poor quality outputs. This problem is particularly acute in specialized fields like ECG, where precise terminology is critical. To address this gap, we train a Byte Pair Encoding (BPE) tokenizer to incorporate ECG-specific vocabulary, resulting in improved coherence and relevance in text generation. The custom tokenizer was integrated into a GPT-2 model for image-to-text generation tasks, where ECG reports were generated from ECG waveforms. Our experiments show that using the custom tokenizer leads to a 85% improvement in coherence and a 45% increase in relevance of generated reports compared to a general-purpose tokenizer.