Long short-term memory-based chemical language models for bioactive molecular generation using tailored pre-training datasets

Ryuto Abe, Tomoyuki Miyao · Artificial Intelligence in the Life Sciences · 2026

Chemical language models (CLMs) computationally generate molecules as line notations, and their usefulness has been demonstrated in various applications. A common approach to producing target-specific molecules is to pre-train a CLM on a large molecular library and then fine-tune it on a dataset of targeted compounds. In this study, we systematically examine how pre-training datasets influence subsequent fine-tuning. For this purpose, six pre-training datasets were prepared, including those with and without structurally liable molecules estimated using RDKit functions, as well as bioactive and non-bioactive molecules assembled from publicly available databases. Six long short-term memory (LSTM)-based CLMs were pre-trained on the six pre-training datasets and then fine-tuned on five target-oriented molecular datasets. We revealed that selecting a pre-training dataset aligned with the researcher's design objective is key to controlling LSTM-based CLMs during fine-tuning, enabling the generation of non-liable and/or selective bioactive molecules without sacrificing diversity, similarity to the fine-tuning dataset, or test-molecule rediscovery.

Read the paper · More papers on PaperTik