Comparative Analysis of Random Forest and Support Vector Classifier for Student Risk Assessment: A Case Study of Colegio de Getafe

Alvin T. Remolado, Deborah G. Brosas · 2023

This research aims to compare the random forest classifier and support vector classifier to the survey data of the risk assessment system for Colegio de Getafe (CDG), a higher educational institution, to identify at-risk students and provide necessary support and interventions. The current manual data collection and analysis is time-consuming and often leads to inaccurate assessments. This study uses machine learning algorithms to analyze student survey data and classify students based on their unique characteristics, such as the differently abled, at-risk, and special needs. This study involves data collection, preprocessing, cleaning, transformation, and modeling using Random Forest and Support Vector Classifier algorithms. Evaluation metrics include accuracy, precision, recall, and F1 score. The results indicate that the Support Vector Classifier has an accuracy of 100%, while the Random Forest Classifier shows an accuracy of 98.09%. The SVC’s ability to capture complex relationships and decision boundaries in the data contributes to its superior performance. The findings highlight the potential of the SVC as a powerful tool for analyzing survey data in higher education and can inform decision-making and targeted interventions by school administrators. Further research can explore the implications of different data mining techniques and parameter configurations on similar datasets.

Read the paper · More papers on PaperTik