Identification of Complicated Lithology with Machine Learning

Liangyu Chen, Lang Hu, Jintao Xin, Qiuyuan Hou, Jianwei Fu, Yonggui Li, Zhi Chen · Applied Sciences · 2025

Lithology identification is one of the most important research areas in petroleum engineering, including reservoir characterization, formation evaluation, and reservoir modeling. Due to the complex structural environment, diverse lithofacies types, and differences in logging data and core data recording standards, there is significant overlap in the logging responses between different lithologies in the second member of the Lucaogou Formation in the Santanghu Basin. Machine learning methods have demonstrated powerful nonlinear capabilities that have a strong advantage in addressing complex nonlinear relationships between data. In this paper, based on felsic content, the lithologies in the study area are classified into four categories from high to low: tuff, dolomitic tuff, tuffaceous dolomite, and dolomite. We also study select logging attributes that are sensitive to lithology, such as natural gamma, acoustic travel time, neutron, and compensated density. Using machine learning methods, XGBoost, random forest, and support vector regression were selected to conduct lithology identification and favorable reservoir prediction in the study. The prediction results show that when trained with 80% of the predictors, the prediction performance of all three models has improved to varying degrees. Among them, Random Forest performed best in predicting felsic content, with an MAE of 0.11, an MSE of 0.020, an RMSE of 0.14, and a R2 of 0.43. XGBoost ranked second, with an MAE of 0.12, an MSE of 0.022, an RMSE of 0.15, and an R2 of 0.42. SVR performed the poorest. By comparing the actual core data with the predicted data, it was found that the results are relatively close to the XRD results, indicating that the prediction accuracy is high.

Read the paper · More papers on PaperTik