Ensemble Model with BERT, RoBERTa and XLNet for Molecular Property Prediction
Junling Hu · ICCK Transactions on Emerging Topics in Artificial Intelligence · 2026
Molecular property prediction is a fundamental task in drug discovery and materials science, yet most high-performing approaches depend on large-scale pretraining that demands substantial computational resources. This work proposes a pretraining-free ensemble framework that trains multiple Transformer-based architectures—BERT, RoBERTa, and XLNet—from random initialization using the Atom-in-SMILES (AIS) molecular representation, which provides richer atomic-level semantics than conventional SMILES. The three Transformer encoders are coupled with BiLSTM prediction heads and integrated via a BaggingRegressor to reduce variance and improve generalization. Experiments on the ZINC250k and ZINC310k benchmarks demonstrate that the proposed framework achieves competitive performance against pretrained baselines including GROVER, CHEM-BERT, and D-MPNN, while requiring only task-specific end-to-end training with adaptive early stopping. These results establish that carefully designed molecular representations combined with heterogeneous ensemble learning can serve as a practical and resource-efficient alternative to pretraining-based paradigms in molecular modeling.