BioReadNet: A Transformer-Driven Hybrid Model for Target Audience-Aware Biomedical Text Readability Assessment
Anya Amel Nait Djoudi, Patrice Bellot, Adrian-Gabriel Chifu · 2025
The perception of the readability of biomedical texts varies depending on the reader's profile, a disparity further amplified by the intrinsic complexity of these documents and the unequal distribution of health literacy within the population. Although 72% of Internet users consult medical information online, a significant proportion have difficulty understanding it. To ensure that texts are accessible to a diverse audience, it is essential to assess readability. However, conventional readability formulas, designed for general texts, do not take this diversity into account, underlining the need to adapt evaluation tools to the specific needs of biomedical texts and the heterogeneity of readers. To address this gap, we propose a novel readability assessment method tailored to three distinct audiences: expert adults, non-expert adults, and children. Our approach is built upon a structured, bilingual biomedical corpus of 20,008 documents (8,854 in French, 11,154 in English), compiled from multiple sources to ensure diversity in both content and audience. Specifically, the French corpus combines texts from Cochrane and Wikipedia/Vikidia, both of which are subsets of the CLEAR corpus, while the English corpus merges documents from the Cochrane Library, Plaba, and Science Journal for Kids. For each original expert-level text, domain specialists produced simplified variants calibrated specifically to the comprehension abilities of non-expert adults or children. Every document is therefore explicitly labeled by its target audience. Leveraging this resource, we trained a diverse suite of classifiers, from classical approaches (e.g., XGBoost, SVM) to classifiers built upon language models (e.g., BERT, CamemBERT, BioBERT, DrBERT). We then designed a hybrid architecture "BioReadNet" that integrates transformer embeddings with expert-driven linguistic features, achieving a macro-averaged F1 score of 0.987.