An Improved Model for Predicting Compound Retrosynthesizability Using Machine Learning
Mami Ozawa, Nobuaki Yasuo, Masakazu Sekijima · 2022
Computational science has attracted significant attention as a means of reducing research and development costs in the field of drug discovery. An example of compound discovery is to propose numerous compounds using molecular generation models, and then filter them by determining whether they can be synthesized or not by a computer before proceeding to wet experiments. Filtering can be done by retrosynthesis or by scoring. An example of existing scoring methods is RAs-core, which predicts the retrosynthetic accessibility by machine learning. Therefore, this study discusses the practicality of the RAscore model and its associated problems. It also proposes a method for building a model that can make more accurate predictions based on hypotheses about the causes of the problem. In addition to the ChEMBL data used to train the original model, we created three models using data created with the ChemTS molecular generation model, and selected the model with the best evaluation results. The models were evaluated by comparing scores using existing RAscore models as a baseline, and one model of the three models achieved a better AUC (area under the curve) and binary accuracy. The models, datasets, and source code are available at https://github.com/sekijima-lab/retrosynthesizability_prediction_models.