Cyber-Risk Forecasting using Machine Learning Models and Generalized Extreme Value Distributions

Jules Sadefo Kamdem, KAPSA, Danielle SELAMBI · HAL (Le Centre pour la Communication Scientifique Directe) · 2022

In this paper, we estimate the cost of a data breach using the number of compromised records. The number of such records is predicted by means of a machine learning model, particularly the Random Forest. We further analyse the fat tail phenomena which capture the underlying dynamics in the number of affected records. The objective is to calculate the maximum loss in order to answer the question of the insurability of cyber risk. Our results show that the total number of affected records follow a Frechet distribution, and we then estimate the Generalized Extreme Value (GEV) parameters to calculate the value at risk (VaR). This analysis is critical because it gives an idea of the maximum loss that can be generated by an enterprise data breach. These results are usable in anticipating the premiums for cyber risk coverage in the insurance markets.

Read the paper · More papers on PaperTik