Dual-target candidate compounds from a transformer chemical language model contain characteristic structural features

Sanjana Srinivasan, Alec Lamens, Jürgen Bajorath · European Journal of Medicinal Chemistry Reports · 2025

Chemical language models (CLMs) are increasingly used for generative design of candidate compounds for medicinal chemistry. However, their predictions are difficult to rationalize. Currently, detailed computational explanations of CLM-based compound generation are unavailable. Therefore, we have attempted to better understand from a medicinal chemistry perspective how CLMs learn and arrive at compound predictions. Therefore, we have subjected dual-target candidate compounds for polypharmacology generated with transformer CLMs to a series of analysis steps exploring structural features that are learned and compared them to known compounds with dual-target activity. Using machine learning combined with distinct chemical structure-oriented approaches from explainable artificial intelligence, we show that CLMs learn substructures characteristic of known dual-target compounds as a basis for generating new candidates with various chemical modifications. A representative compound with dual-target activity (center) is surrounded by candidate compounds generated using a chemical language model. A shared substructure distinguishing these molecules from corresponding single-target compounds is highlighted in pink. • Exploring learning characteristics of chemical language models. • Medicinal chemistry-centric analysis of dual-target compounds. • Analysis scheme combining machine learning and XAI methods. • Identification of characteristic substructures driving predictions. • Chemically intuitive explanations of generative compound design.

Read the paper · More papers on PaperTik