Predicting Chronic Kidney Disease in Type 2 Diabetes Using Natural Language Processing on Healthcare Data
Juan F. Navarro‐González, Leopoldo Pérez de Isla, Gloria Cánovas Molina, Miguel Ángel Brito-Sanfiel, David E. Barajas Galindo, Luís Ángel Cuellar Olmedo, Dı́dac Mauricio, Santiago Tofé, José Antonio Balsa Barro, Matilde Rubio-Almanza, José Juan Aparicio Sánchez, Miren Sequera Mutiozabal, Belén Pimentel, Ana Pérez Domínguez, Carlos Arias-Cabrales, Víctor Fanjul, Antonio Jesús Blanco-Carrasco, Juan Francisco Merino-Torres · Kidney Diseases · 2025
Introduction: Persons with type 2 diabetes mellitus (T2DM) attending hospitals frequently experience major complications. We assessed the potential use of unstructured free-text data extracted from electronic health records (EHRs) using natural language processing (NLP) and machine learning (ML) to develop a predictive model for chronic kidney disease (CKD) in T2DM. Methods: This multicenter retrospective study included data from eight Spanish hospitals (2013-2018), extracted using NLP and ML techniques (EHRead®) based on SNOMED CT terminology. From a cohort of individuals with T2DM, we identified those with and without CKD at inclusion. Among individuals without CKD, we trained and validated a 2-year predictive model for CKD development. The model showing the best balance between performance and clinical interpretability was selected for integration into a web-based tool to support early detection and risk stratification. Results: Of 588,786 individuals with T2DM, 316,597 were included for model development (training: 291,429 [92.1%]; validation: 25,168 [7.9%]; CKD incidence: 15.4% and 18.4%, respectively). A high proportion of missing data was observed in key clinical variables. Among models evaluated, logistic regression achieved the best performance (receiver operating characteristic area under the curve 0.72) using 27 predictors. Both a reduced 10-predictor model and a clinically refined 8-predictor model showed comparable performance to the full model in training and validation cohorts. The clinically refined model was selected for implementation in the web-based tool. Conclusion: Unstructured EHR data enabled the development of a predictive model for 2-year CKD risk in persons with T2DM. Improving EHR data completeness remains essential to enhance future predictive modeling.