Developing Best Practices for Descriptor‐Based Property Prediction: Appropriate Matching of Datasets, Descriptors, Methods, and Expectations
Michael P. Krein, Tao-wei Huang, Lisa Morkowchuk, Dimitris K. Agrafiotis, Curt M. Breneman · 2012
After undergoing decades of development and nearly continuous use (and abuse) in a number of application areas, quantitative structure–activity relationship (QSAR) and related descriptor-based statistical learning methods have earned mixed reviews throughout their checkered past. Why is this? To find answers to this question, it would seem that one needs only to examine the capabilities and limitations of each component of modern QSAR/quantitative structure–property relationship (QSPR) methods in some detail and then define the scope and applicability of each one. While this approach may appear useful on the surface, in practice it is necessary to select not only the specific combinations of descriptors and machine learning methods that might work best, but also consider the nature and size of the training and test datasets and the physical effects that could control the endpoints to be modeled – this is where the definition and application of “best practices” is important for obtaining actionable modeling outcomes, and for setting user expectations of modeling accuracy when predicting the endpoint values of unknowns. This chapter explores a wide variety of statistical learning approaches, descriptor types, and model validation strategies and seeks to help end users understand the factors involved in creating and using QSAR/QSPR models effectively.