ORFPPrediction: Machine Learning Based Online Recruitment Fraud Probability Prediction

Haiyan Zhang, Meifeng Wang, Yuanyuan Wang, Yazhou Li, Dingyun Gu, Yemeng Zhu · 2023

In response to the issue of inadequate public datasets concerning imbalances in Internet recruitment scams, this paper constructs an imbalanced experimental dataset, labels it, and proposes an ensemble learning method based on model stacking. This method enhances the prediction accuracy of the original model built with Random Forest. Initially, the method employs label encoding and TF-IDF (Term Frequency- Inverse Document Frequency) to preprocess structured and unstructured data respectively. Then, it utilizes Truncated SVD (Singular Value Decomposition) to reduce the dimensionality of the sparse matrix generated by feature extraction from TF-IDF. Subsequently, decision trees, random forests, and Light Gradient Boosting Machine are utilized as base models, with the random forest employed as the meta-model to integrate the predictions of the base models. This paper presents a comparison of the experimental results from seven machine learning models based on the self-constructed experimental dataset, with model stacking implemented on the random forest that yielded the best experimental result. Through model stacking enhancement of the machine learning models, compared to the random forest model that performed best on the self-constructed experimental dataset, an accuracy of 94.85% was achieved, marking a 7.33% improvement. The experimental results hold significant value and practical implications in understanding and resolving the probability risk issue of recruitment fraud.

Read the paper · More papers on PaperTik