Exploratory Methods for Imbalanced Data Classification in Online Recruitment Fraud Detection: A Comparative Analysis

Jiaxu Li, Yunxuan Li, Hongjian Han, Xin Biao Lu · 2021

Online recruitment platforms have been increasingly used by companies and applicants. However, there have been a growing number of online recruitment frauds (ORFs) in recent years, seriously affecting the company's reputation and accounting for personal information leakage. On the other hand, identifying ORF with classification models is challenging, as the ORF datasets are typically highly imbalanced such that the accuracy in predicting the minority class in practical recruitment systems is not satisfiable. To overcome these limitations, with empirical Employment Scam Aegean Dataset (EMSCAD), we implement a comprehensive comparative evaluation on processing approaches for imbalanced data, including data sampling techniques, cost-sensitive learning, and ensemble learning. And we design a LightGBM ORF detection model based on hybrid sampling. The results indicate that our model has a higher value in F1-measure, precision and recall of 0.93, 0.93 and 0.94, respectively, and that comparative analysis shows that the model with a combination of data sampling and ensemble learning can achieve improved performance in finding frauds in ORF datasets.

Read the paper · More papers on PaperTik