House price prediction using clustering and genetic programming along with conducting a comparative study
Fateme Azimlu, Shahryar Rahnamayan, Masoud Makrehchi · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2021
One of the most important tasks in machine learning is prediction. Data scientists use different regression methods to find the most appropriate and accurate model for each type of datasets. This study proposes a method to improve accuracy in regression and prediction. In common methods, different models are applied to the whole data to find the best model with higher accuracy. In our proposed approach, first, we cluster data using different methods such as K-means, DBSCAN, and agglomerative hierarchical clustering algorithms. Then, for each clustering method and for each generated cluster we apply various regression models including linear and polynomial regressions, SVR, neural network, and symbolic regression in order to find the most accurate model and study the genetic programming potential in improving the prediction accuracy. This model is a combination of clustering and regression. After clustering, the number of samples in each created cluster, compared to the number of samples in the whole dataset is reduced, and consequently by decreasing the number of samples in each group, we lose accuracy. On the other hand, specifying data and setting similar samples in one group enhances the accuracy and decreases the computational cost. As a case study, we used real estate data with 20 features to improve house price estimation; however, this approach is applicable to other large datasets.