Bayesian transfer learning with Monte Carlo Markov Chains for kinetic modelling of pilot plant and industrial data
Per Julian Becker, Warumporn Pejpichestakul, Benoît Celse · Digital Chemical Engineering · 2026
This work presents a methodology for kinetic model parameter fitting with Transfer Learning to improve model robustness and mitigate over-fitting of poorly sensitized parameters. Datasets from lab, pilot plant, or industrial scale experiments for the same process often cover different regions of the design space. Therefore, some of the model’s features lack sufficient variations to obtain a robust estimate of the model parameters from one of the datasets alone. Combining such, heterogeneous, datasets is a challenging task and requires extensive expert knowledge. Transfer Learning with Monte Carlo Markov Chains (MCMC) is used to retain information from different datasets. This method uses the Bayes theorem to impose a prior distribution on the model parameters when sampling the likelihood distribution via repeated model inference for random parameter variation. The aim of the methodology is to leverage the respective strengths and mitigate the weaknesses of different datasets, which will lead to a more robust estimation of the kinetic model parameters. In this work we use the MCMC algorithm for parameter identification for a hydrodenitrogenation (HDN) model used to simulate the pretreatment reactor of hydrocracking units. Industrial follow-up and pilot plant data was used in this study. A deactivation model was developed to account for loss of catalyst activity during long cycles of industrial units. The kinetic parameters are well sensitized in the pilot plant dataset, but feed variability is low. The industrial dataset has more variation of feedstock descriptors but temperature, pressure, and LHSV effects are poorly sensitized. The comparison to a naive approach of identifying selected subsets of the parameters on each dataset shows that the Transfer Learning method provides less over-fitting by retaining information from both datasets.