Methodology and Evaluation for Machine Learning in Drug Molecule Design
Zhiyuan Yang · Springer Link (Chiba Institute of Technology) · 2026
With the aging of world populations and the increasingly complex disease conditions, traditional drug discovery also encounters some problems such as soaring costs, extended time frames, and slim winning odds. The enormous chemical space has made laboratory screening more difficult to do. And further tangling constraints, like optimizing pharmaceutical properties and synthesizability, make it more difficult to identify lead compounds. The advances in biological data collection and high-performance computing lead this review on the topic of generative deep learning in drug molecule design. This review delineates the capabilities and limitations of linear and graph-based representations and summarizes the operating principles and strengths of commonly used machine learning architectures. In addition, this review presents a twofold evaluation process, i.e., generation performance and drug-likeness. Across recent studies, the generative approaches speed up drug-like molecule exploring and enable goal-conditioned design, although faithful 3D conformations and integration with molecular dynamics still have some problems, and a sound quantification of synthetic accessibility still is a problem. It’s possible that a pipeline that is closed-loop and couples the generation of molecules with retrosynthesis and automated experimentation would be advantageous.