Healthcare Simulation Data Generation Program
Song Yu, Masaki Matsumori, Yizhen Bao, Kenji Baba, Liangyi Sun, Fenglai Qin · 2022
The lack of training data for models has always been an important issue in the construction of machine learning models. In this paper, a simulation data generation procedure was built to obtain more healthcare data with a small amount of sample. This simulation data belongs to healthcare data (64 items in total), which specifically includes: basic personal information items, medical expenses items, common diseases items, daily living habits items and health checkup values items. In addition to the basic personal information items, the data is divided into continuous data items (medical expenses and health check-up values) and discrete data items (common diseases and daily habits). Based on the continuous data, a large amount of continuous data is generated from its multivariate distribution, while ensuring that the distribution satisfies the Central Limit Theorem. A random forest regression model is built to simulate the medical cost data. As for the discrete data, a Support Vector Machine (SVM) was built and then an optimized SVM model was obtained by multivariate comparison and by considering the loss function, penalty term and regularization. To validate the model soundness, the model fit of the random forest regression model was tested using a measure of the relationship between prediction error and mean expectation, R^2. An accuracy assessment model for the simulation was established to assess the prediction efficiency of the optimized SVM model, which was assessed to be essentially above 80%. In summary, the simulation generation procedure provides technical support to improve the medical database.