Leveraging ML to Detect Fraudulent Customs Declaration: A Case of Logistic Regression and LightGBM Algorithms

Innocent Chilomo, Bennett Kankuzi · 2024

Harmonised System (HS) code fraud, also known as Customs tariff classification fraud, is common in developed and developing countries. It occurs when traders intentionally mis-classify goods during the time of clearance. Classification fraud results in tax revenue loss, restricted and prohibited goods entry, and unfair market competition. Current manual verification techniques to detect classification fraud are time-consuming and error-prone; thus, an effective and precise strategy is required. In this study, we developed and evaluated the performance of Logistic Regression and LightGBM machine learning models in detecting HS Code fraud using a dataset obtained from a government Revenue Authority in a developing country. Results show that the LightGBM model outperformed Logistic Regression. The LightGBM model demonstrated superior accuracy (0.961), precision (0.890), recall (0.901), and F1 score (0.896) as compared to the Logistic Regression model, which had accuracy (0.932), precision (0.772), recall (0.900), and F1 score (0.831). Our contribution is threefold: Firstly, the dataset used is from a developing country, unlike other studies that use datasets from developed countries. Secondly, we worked with Customs import data collected at an 8-digit HS code level, whereas most similar studies typically used HS codes at 2, 4, or 6-digit levels. Lastly, unlike comparable studies, we have compared the performance of Logistic Regression and LightGBM machine learning algorithms.

Read the paper · More papers on PaperTik