Evaluation Datasets for GPT-3.5 and GPT-4 in Molecular Generation Tasks

Jinlu · Figshare · 2023

This collection includes datasets from ChEBI-20 and ZINC utilized for the performance evaluation of GPT-3.5 and GPT-4 models in three distinct tasks: Molecular Description Generation, Molecular Optimization, and Molecular Generation. The ChEBI-20 dataset was used for the Molecular Description Generation and Molecular Generation tasks, where the LLMs were required to generate accurate descriptions of molecular functional groups and viable molecular structures from scratch, respectively. The ZINC dataset was employed for the Molecular Optimization task, which involves the modification of molecules to optimize a given property, such as the Quantitative Estimate of Drug-likeness (QED). For each task, we selected 10 cases for testing.

Read the paper · More papers on PaperTik