Comparing Regression Models Predicting the Price of Used Cars in Big Data

Jo In Kang, Heta Parekh, Priya Ramdas, Seongwon Lee, Jongwook Woo · 2022 IEEE International Conference on Consumer Electronics-Asia (ICCE-Asia) · 2022

We built and compared regression models to find the optimal solution using Spark Big Data engine, which predicts the price of used cars. We have the Cargurus dataset to perform predictive analysis, which is consumer data set as of 9.29 GB. The traditional systems have difficulty to handle the large scale data greater than hundreds of Mega-Bytes. Our Spark cluster can resolve this issue by storing and training the models with the large-scale data as distributed parallel computing systems. We adopt multiple regression algorithms to train models using Spark library in the big data environment. Then, we compare the models by measuring the computing time and accuracy. We observed that the GBT model shows the best accuracy in RMSE and R2 but relatively long computation time. The Random Forest model has a similar RMSE and R2 as GBT but with less computing time.

Read the paper · More papers on PaperTik