An analysis of the proportion of feature subsampling on XGBoost - A case study of claim prediction in car insurance

Wafiyatul Khusna, Hendri Murfi · AIP conference proceedings · 2020

Claim prediction is one of the important elements in the insurance. The increasing frequency of claim makes the data volume also increases to become big data. So, we need the right machine learning method to help insurance companies manage big data more efficiently. XGBoost is a machine learning model based on decision trees. XGBoost can be applied for claim prediction case in the form of two-class or multi-class classification. We may select a subset of features in building the XGBoost model especially for data with a large number of features. In this paper, we examine the influence of the proportion of features on the accuracy of the XGBoost model. Our simulations show that by randomly using 1/5 of features, the XGBoost model can produce accuracy comparable to the model that uses all features. It means that the XGBoost model is scalable in terms of the proportion of features.

Read the paper · More papers on PaperTik