Enhancing Drug Sensitivity Prediction in Cancer Cell Lines Using Multi-Omics Data and Machine Learning

Sara Amjad, Mehar Aziz, M. M. Sufyan Beg · 2025

This study integrates diverse omics datasets, including proteomic, transcriptomic, and miRNA profiles from cancer cell lines, with drug sensitivity (AUC) data for 285 drugs. We identified that combining proteomics, transcriptomics, and miRNA data provides the best predictive model. Using the k-means clustering algorithm, we grouped cell lines based on these multi-omics data. Subsequently, drug sensitivity (AUC) for 285 drugs was predicted using a 5-fold cross-validation Random Forest Classifier for each cluster separately. The resulting five clusters were thoroughly evaluated, leading to drug suggestions with over 75% accuracy for further investigation. Our findings indicate that high-quality clusters derived from the optimal set of multi-omics features improve the performance of the Random Forest algorithm in predicting drug responses, compared to low-quality clusters obtained from a random set of features. We also compared drug similarities within clusters 0, 1, and 2, analyzing the results with a heatmap. Our findings show improved similarity between drugs in cluster 0 compared to clusters 1 and 2, based on hydrogen bond donor, hydrogen bond acceptor, and molecular weight properties. This aligns with the silhouette scores, which is higher for cluster 0 compared to clusters 1 and 2. These results suggest that, in high-quality clusters, drugs with high prediction accuracy tend to exhibit greater similarity in their properties compared to those in low-quality clusters. This supports a cluster-centric approach in precision oncology, as high-quality clusters, indicated by high silhouette scores, lead to more accurate and consistent drug response predictions.

Read the paper · More papers on PaperTik