Optimizing clustering of electronic health tabular data: generative adversarial networks and Dirichlet process mixture models for advance healthcare analytics

Francis John Kita, Gadde Srinivasa Rao, Peter Josephat Kirigiti · IISE Transactions on Healthcare Systems Engineering · 2025

The use of electronic health record (EHR) tabular data for clustering has garnered global attention for its potential to uncover meaningful clusters, improve patient care and advance healthcare research. Applications of EHR data extend beyond healthcare to include financial modeling, fraud detection, differential diagnosis and rare disease identification. However, critical challenges persist, including privacy concerns, limited access to real EHR data, underexplored generative methods and inadequate validation frameworks. This study aims to generate and validate synthetic EHR tabular data and demonstrate its utility in clustering, addressing existing gaps in privacy, data quality and validation. Advanced deep learning architectures, including Generative Adversarial Networks (GANs), were developed to generate high-quality synthetic datasets. We implemented robust validation frameworks that incorporate fidelity, utility and privacy-preserving mechanisms. Dirichlet Process Mixture Models (DPMMs) used synthetic data to group similar items and measured the differences between real and synthetic data distributions by estimating divergence. GANs effectively replicated complex datasets while balancing fidelity, utility and privacy. Validation confirmed the robustness of DPMMs in identifying latent structures. Cluster analysis revealed significant multimorbidity patterns, facilitating personalized treatment strategies, resource allocation and preventive care. This study demonstrates the transformative potential of synthetic EHR data in clustering, advancing healthcare delivery, planning and research.

Read the paper · More papers on PaperTik