Correlation network construction for molecular data based on cross validation framework
Prabhakar Chalise, Indrani Sarker, Yanming Li · Computers in Biology and Medicine · 2026
Correlation networks are widely used to uncover latent interaction structures within biological datasets. These networks are typically constructed using Pearson or partial correlation coefficients. Pearson correlation assesses pairwise associations without accounting for other variables, whereas partial correlation measures associations while adjusting for potential confounders. In both approaches, network construction relies on statistical significance testing. Because statistical significance depends heavily on sample size, these methods may result in under- or over-estimate the network structure failing to accurately capture the underlying biological interactions. To address this limitation, we propose a novel cross-validation-based method, cvCorNet, for constructing biological correlation networks. The dataset is partitioned into training and test sets. Nodewise regression models are fitted on the training data, and partial correlation networks are predicted for the test data using a user-specified cutoff value. Simultaneously, partial correlation networks are independently estimated from the test data. The optimal cutoff is then selected by maximizing the agreement between the predicted and empirically estimated networks across a range of correlation thresholds. The performance of cvCorNet is evaluated under a broad spectrum of simulated scenarios. Results demonstrate that cvCorNet effectively recovers the underlying correlation structure, particularly by controlling false positive edges. The application of the method is further illustrated using two real-world datasets from glycomics and acute myeloid leukemia studies.