EHealth: A Chinese Biomedical Language Model Built via Multi-Level Text Discrimination
Quan Wang, Songtai Dai, Benfeng Xu, Yajuan Lyu, Hua Jian Wu, Haifeng Wang · IEEE Transactions on Audio Speech and Language Processing · 2025
Pre-trained language models (PLMs) have recently revolutionized the field of natural language processing, impacting not only the general domain but also the biomedical domain. Most previous studies on constructing biomedical PLMs relied simply on domain adaptation and focused mainly on English. This work introduces eHealth, a compact, encoder-based Chinese biomedical PLM that can be fine-tuned in a customized manner to effectively handle various Chinese biomedical language understanding tasks. Rather than relying on domain adaptation, eHealth is built from scratch using a novel pre-training framework. This framework trains eHealth as a discriminator through token- and sequence-level discrimination. Token-level discrimination detects corrupted input tokens and recovers their original forms from plausible candidates, while sequence-level discrimination further distinguishes corruptions of the same original sequence from those of others. As a result, eHealth can learn language semantics at both token and sequence levels. Moreover, eHealth is trained entirely from scratch with a newly constructed in-domain vocabulary, which may help improve tokenization and understanding of biomedical text. Extensive experiments on 11 Chinese biomedical language understanding tasks of various forms verify the effectiveness and superiority of eHealth. The pre-trained model as well as the code have been released to the public.