FINE-TUNING CHEMBERTA TRANSFORMER MODEL FOR SUPERIOR MOLECULAR PROPERTY PREDICTION (MPP)
Journal of Theoretical and Applied Information Technology · Journal of Theoretical and Applied Information Technology · 2025
Drug discovery is expensive and time-consuming, often taking decades and billions of dollars for its research and development to analyze its chemical properties for drug development. The task of Molecular Property Prediction (MPP) is major crucial activity for selecting potential drug candidates in drug Discovery. The chemical language models (CLMs) have gained attention in this area due to their success in natural language processing tasks like text analysis and translation. The availability of vast unlabeled chemical data has driven the development of CLMs to interpret molecular information effectively. This study uses a large language model approach by fine-tuning a Transformer model to predict molecular properties such as solubility, toxicity and pIC50. Traditionally, molecular graphs and descriptors were used for property prediction. However, Transformer-based models like BERT, GPT, and their variations have shown exceptional performance in downstream NLP tasks. In this research, the chemical informatics, model such as ChemBERTa have been explored for learning molecular contextual information from Simplified Molecular Input Line Entry System (SMILES) sequences. Experimental results show that ChemBERTa outperforms traditional ML models in molecular property prediction, achieving Accuracy and AuROC scores of 0.94 and 0.96, respectively. This demonstrates its superior predictive capability compared to state-of-the-art methods. The study highlights the potential of Transformer-based CLMs in accelerating drug discovery by effectively predicting molecular properties from text-based inputs like SMILES strings