A Novel Random Forest Classifier with K-Fold cross validation for crime prediction using the San Francisco dataset and comparison with Stochastic Gradient Descent classifier for improved accuracy

Zeenath Rahaman, R Mahaveerakannan · 2024

This study uses the San Francisco dataset and applies the Novel Random Forest Classifier with K-Fold Cross Validation to improve the accuracy of crime prediction. The outcomes derived from this methodology will be contrasted with those produced by the stochastic gradient descent classifier. The procedure entails forecasting criminal activities by utilizing the Novel Random Forest Classifier with K-Fold Cross Validation and Stochastic Gradient Descent. This classification algorithm utilizes a range of training and testing split settings. In addition, we choose 4,763 individuals from each of the two categories for study. The Gpower test, conducted with a significance level (α) of 0.05 and a power setting of 0.80, produces results with a confidence level of 80%. The Novel Random Forest Classifier outperforms the Stochastic Gradient Descent Classifier, which attained an accuracy of 94.7320 percent. The Independent sample t-test yielded a significant p-value of 0.01 (p<0.05), indicating that the study comparing the Novel Random Forest Classifier with the Stochastic Gradient Classifier has a statistically significant outcome. Hence, it can be inferred that the Novel Random Forest Classifier with K-Fold Cross Validation surpasses the Stochastic Gradient Descent Classifier in crime prediction, as indicated by its superior accuracy %.

Read the paper · More papers on PaperTik