Machine Learning-based Adverse Drug Reaction Prediction: Model Comparisons, Feature Optimization and Generative AI Challenges

Ana Maria Cucos, László Barna Iantovics · Procedia Computer Science · 2025

Adverse drug reactions (ADRs) pose significant challenges in pharmacovigilance, necessitating predictive models for early detection. In this study, we leverage the SIDER database, comprising 28 attributes, where the first attribute contains SMILES strings representing chemical structures, and the remaining 27 attributes correspond to binary-labeled system organ classes for ADRs. We employ three machine learning algorithms—Support Vector Machine (SVM), Logistic Regression (LR), and AdaBoost—to classify ADR occurrences based on molecular features. LR was chosen for its interpretability and baseline performance, SVM was chosen for its effectiveness in handling imbalanced datasets and AdaBoost was chosen for its ability to improve weak learners and handle class imbalance through adaptive weighting. Three distinct feature sets are generated from SMILES strings: (1) molecular fingerprints (Morgan, Atom Pair, Torsion, and Pattern), which are known to capture local environments (Morgan), atom relationships (Atom Pair), conformational flexibility (Torsion) and substructure motifs (Pattern), (2) molecular descriptors (Molecular Weight and Number of Valence Electrons), because they influence molecular size, reactivity and overall chemical behavior, and (3) a hybrid combination of both fingerprints and descriptors. To optimize model performance, we apply Random Search, which randomly samples hyperparameter values within specified ranges to find optimal model settings efficiency, for hyperparameter tuning and compare the predictive efficacy of each algorithm across different feature sets. To compare the predictive efficacy, Accuracy, AUC-ROC and AUC-PR metrics were used to evaluate overall correctness, discrimination ability and performance on imbalance data. Our findings provide insights into the suitability of diverse molecular representations for ADR prediction (pADR) and highlight the impact of algorithmic selection and hyperparameter optimization on classification performance. The results demonstrate that SVM outperformed LR and AdaBoost in predicting ADRs, especially when using Morgan fingerprints, achieving an AUC-PR of 0.75 for Metabolism and nutrition disorders ADR. Descriptors-based models showed lower performance, but combining fingerprints and descriptors led to slight accuracy improvements. A major challenge in generative AI for ADR prediction is creating realistic synthetic data while avoiding bias and maintain clinical relevance. Pharmacists and clinical researchers can benefit from ADRs predictions to improve patient safety, optimize prescriptions and design safer treatments.

Read the paper · More papers on PaperTik