Exploration of read-across-based predictions for unveiling the risk of FDA-labeled drug-induced cardiotoxicity (DICT) rank: A careful classification of the most extensive reference list of human drugs
Sapna Kumari Pandey, Kunal Roy · In Silico Research in Biomedicine · 2025
• We develop machine learning (ML) models to accurately classify the DICTrank dataset of the US-FDA. • We use structural similarity features along with standard molecular features in developing multiple ML models. • The developed similarity-based ML models with a lower number of features show acceptable external validation results. Drug-induced cardiotoxicity (DICT) is considered one of the primary reasons for drug attrition and the core issue in non-clinical and clinical safety testing of new and existing drugs. In this work, we investigated the impact of standard molecular features and structure-derived similarity features in advanced machine learning (ML) algorithms for accurately classifying the DICTrank dataset provided by the United States Food and Drug Administration (US FDA). Several similarity-derived read-across (RA) approaches, including structural similarity, biological similarity, and similarity based on mode of action, are widely accepted in the regulatory context. This study investigates the utility of incorporating structural similarity features derived from RA alongside standard molecular features in developing multiple ML models. Based on the acquired results, we found that the number of features included in similarity-derived ML models is lower than in molecular feature-derived models, although they show comparable validation outcomes. The developed models give acceptable external validation results (Matthews correlation coefficient (MCC) = 0.105–0.553, Cohen's kappa 0.205–0.547), justifying the reliability and usefulness of our models. Our research aims to gain a deeper understanding of chemical categorization by combining insights from structural similarity features with molecular features, enabling more rational predictions of query drug molecules.