Advancing chemical grouping: development and application of signature-based structure-activity groups for non-animal safety assessments
J Muldoon, H. Moustakas, Terry Wayne Schultz, T.M. Penning, Amanda Bryant-Friedrich, D. Botelho, A.M. Api · Computational Toxicology · 2025
The Research Institute for Fragrance Materials, Inc. (RIFM) has developed a robust, reliable, reproducible method for clustering chemicals based on their structural signatures and deriving structure–activity groups. This method facilitates the institutionalization of knowledge gained from manually assessing thousands of chemical pairings of fragrance ingredients. The technique improves accuracy, consistency, transparency, and explainability for evaluating chemical safety while reducing reliance on expert judgment and any associated bias. A material’s signature-based structure–activity group is created via a top-down approach using standardized signature trees based on Indicator Phrases (IPs) representing seminal sub-structural features. We have applied the approach to over 6,000 discrete fragrances and fragrance-like organic chemicals (e.g. organic compounds of the chemical classes described in the inventory such as aldehyde, ketone, esters, etc.), and it has been shown to perform well for various properties and parameters observed in this chemical space. The signature trees are adaptable and can be expanded for IPs not found in fragrance materials. The structure–activity groups readily allow for transparent and repeatable separation of an inventory of thousands of chemicals into clusters of chemicals that share the same IPs. Adjacent groups that share all but one or two of the same IPs can be identified, thereby effortlessly expanding the range of potential read-across source substances. With its ease of interpretation, the system facilitates discussions among scientists with different levels of chemical knowledge. In addition to clustering for data-gap filling through read-across, other applications include prioritization for testing and predictive toxicology by encoding IPs using various machine-learning techniques.