On the replicability of corpus-derived medical word lists

Cosmin Mihail Florescu, Ryosuke L. Ohniwa · Applied Corpus Linguistics · 2025

• Replicability of findings can be used to develop medical vocabulary lists. • Visual mapping of keyness and dispersion data can inform threshold value setting. • Keyness is an effective measure in identifying domain-specific vocabulary. • Even dispersion across corpus parts may indicate common words. • Vocabulary profilers can provide a rough approximate of word list specificity. Several English medical vocabulary lists have been developed using corpora compiled from a variety of medical texts including research articles and medical textbooks. List items have been identified for inclusion using criteria mostly adopted from previous studies focused on academic vocabulary. This study aims to employ a systematic approach in compiling a corpus to create a medical word list for learners of English aiming to study or practice medicine in an English-speaking country. A large corpus of medical textbooks (CoMeT; 28,384,681 running words) was created using SketchEngine and analyzed to extract high-frequency lemmas. Keyness and dispersion values for each lemma were plotted in a histogram to visualize clustering patterns. This visual map was used to determine threshold values separating a medical vocabulary subset from a general vocabulary subset. The replicability of the findings was evaluated using two corpora (one medical, one non-medical) different from CoMeT. The newly developed list (Core Medical List; CoMeL) comprising a total of 2881 lemmas was found to include significantly more medicine-specific words and to have higher replicability compared to existing lists. CoMeL may assist learners and educators in English for Medical Purposes programs, including those aiming to undertake challenging medical licensing examinations in English-speaking countries.

Read the paper · More papers on PaperTik