Fairer benchmark of group contribution and machine learning models for property prediction: A new data splitting strategy

Adem R.N. Aouichaoui, Jingkang Liang, Jens Abildskov, Gürkan Sin · Computers & Chemical Engineering · 2025

Accurate prediction of thermophysical properties is important in chemical engineering, where group-contribution models (GCM) have been used extensively. The norm when developing GCMs is to use all available data for parameter estimation, preventing a fair comparison with machine learning (ML) methods that require separate training, validation, and testing data. In this study, we first highlight the detrimental effect of missing groups resulting from using conventional split methods (random and cluster-based) in the development of GCM and ML models using groups as features. This was illustrated by developing property models for the critical point properties (critical temperature, critical pressure and critical volume) as well as the acentric factor. To alleviate this, we propose a novel hybrid splitting algorithm that first ensures that all available groups are represented using the smallest subset of compounds possible and then supplements the subset with molecules based on the Butina clustering of the remaining compounds. The methods show performance close to the optimal possible result produced from using all data for the model calibration, and provide a more fair basis for comparing GCM with ML-based methods. We further benchmark the GCM with seven ML techniques (random forest, decision tree, gradient boosting, extreme gradient boosting, Gaussian processes and support vector machines as well as deep neural networks) using groups as features and three graph neural network models (attentiveFP, MEGNet and GroupGAT). The results show that GroupGAT consistently outperforms other methods on the external test dataset, achieving lower errors than both traditional and ML-enhanced GC methods.

Read the paper · More papers on PaperTik