De novo Drug Design – Ye olde Scoring Problem Revisited
Gisbert Schneider, Kimito Funatsu, Yasushi Okuno, David Alan Winkler · Molecular Informatics · 2017
Generating molecules by computational means is common practice in drug discovery. New chemical entities with the desired properties may serve both as tool compounds for chemogenomics studies and as starting points for hit-to-lead expansion. The three challenges for automated de novo design are i) the assembly of synthetically accessible structures, ii) scoring and property prediction, and finally iii) the systematic optimization of promising molecules in adaptive learning cycles.1 During the past three decades, numerous methods, algorithms and heuristics have been proposed to address each of these problems.2 While the generation of new chemical entities with attractive chemical scaffolds has become feasible by reaction-driven fragment assembly, and the in silico optimization problem may also be considered largely solved, the persisting issue of compound scoring remains difficult. Scoring means picking the best compounds from a large pool of accessible possibilities.3 This process typically includes both ligand- and structure(receptor)-based virtual screening of the computationally generated molecules. This large virtual compound pool contains many more inactive or problematic chemical structures than desirable ones. While compound elimination by appropriate scoring models discards the bulk of the designs (“negative design”) with acceptable accuracy, the selection of the best or most promising ones (“positive design”) remains error-prone. The conventional techniques employed at this step of the selection process include coarse-grained and application-specific heuristics,4,5 physicochemical property calculation,6 quantitative and qualitative structure-activity relationship models,7 similarity calculations on various levels of detail and with different molecular representations,8 shape matching and automated ligand docking, as well as the detection of potentially toxic and otherwise unwanted chemical structures.9 More recently, qualitative and quantitative on- and off-target prediction methods have been added to the molecular designer's tool chest.10 During the PacifiChem2015 conference in Honolulu, HI, USA, we organized a symposium to share experience and discuss the progress in computer-based de novo compound design and scoring (Figure 1). Several symposium papers and closely related contributions are compiled in this focused special issue of Molecular Informatics. It is evident that project-specific, customized scoring functions will help reduce the false-positive prediction rate. In this special issue, we highlight methods that diverge from mainstream approaches and point towards future developments in molecular informatics research. In a Methods Corner article, Kaneko and Funatsu review a structure generation method that can be used to design molecules that lie within the applicability domain of a scoring function. This “inverse QSAR” approach to de novo design is now awaiting practical application. Fukunishi et al. present methods to improve docking scores by regression-based correction for use in structure-based virtual screening and molecule design. In a cross-validation study, the weighted scoring functions showed improved accuracy over simpler methods. Multidimensional compound optimization by Pareto ranking is showcased in the article by Daeyaert and Deem. They show that Pareto sorting improves the performance of de novo design algorithms in generating molecules with hard-to-meet constraints. This tackles the important problem of false positive hits in structure based drug design by introducing structural and physicochemical constraints in the designed molecules, and by forcing known essential interactions between these molecules and their target receptor to occur. The issue of large-scale on- and off-target prediction is addressed in two contributions. Hamanaka et al. present very large deep neural network models which they trained to predicting ligand-protein relationships. This study nicely exemplifies the usefulness of deep learning methods for hit and lead discovery.11 Le and Winkler dissected the promise deep learning and other machine learning methods further, by benchmarking deep and shallow neural network performance and describing their applications in QSAR and drug discovery. A new hybrid deep network architecture is presented in the article by Schneider et al., who demonstrate the prospective applicability of their approach by designing new antimicrobial peptides. Button et al. show that macromolecular target panel prediction by fast ligand scoring can actually be applied to de novo generated compounds, and helps prioritize virtual screening hits. Finally, Grissoni et al. introduce matrix-based “holistic” molecular representations as descriptors for similarity searching. They demonstrate the applicability of their approach by retrieving a novel cyclooxygenase 2 inhibitor. Researchers from across the globe discussed the status and future possibilities of de novo drug design in a special symposium held at PacifiChem2015 in Honolulu, Hawaii. We hope that the readers will enjoy this compilation of papers and find inspiration for their own de novo molecular designs.