Explainable AI With Imbalanced Learning Strategies for Blockchain Transaction Fraud Detection
Ahmed Abbas Jasim Al-Hchaimi, M. A. Khalifa, Walid El‐Shafai · Engineering Reports · 2026
ABSTRACT Blockchain networks now support billions of dollars in daily transactions, making reliable and transparent fraud detection essential for maintaining user trust and financial stability. Yet, real‐world blockchain datasets are extremely imbalanced, with fraudulent activity representing less than 1% of all transactions. This imbalance causes conventional machine learning models to achieve deceptively high accuracy while still failing to detect a substantial portion of fraudulent events. To address this challenge, this study evaluates the performance and explainability of three models‐XGBoost, LightGBM, and Decision Tree‐on the Ethereum‐based fraud detection data, in which 58% of transactions are identified as fraud. The methodology combines vast feature engineering, k‐fold cross‐validation, and assorted resampling approaches, such as Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling Nearest Neighbor (ADASYN), to revise the effect of class mismatch. Accuracy, AUC, recall, precision, F1‐Score, and Matthews Correlation Coefficient(MCC) are used to measure model performance, and SHapley Additive exPlanations (SHAP) is utilized to give global and local interpretability. Experimental results show that XGBoost combined with SMOTE or ADASYN yields the strongest performance, achieving a recall over 99%, an AUC of 1.000, and a substantially improved MCC compared to training on the raw imbalanced data. LightGBM presents a favourable precision‐recall balance, and Decision Trees demonstrate significant gains after resampling, despite their simplicity. SHAP analysis reveals that log‐transformed transaction amount, merchant‐based encoding, geographic encoding, and temporal features are the primary contributors to fraud risk. These results are important in highlighting two implications: (i) the importance of dealing with extreme class imbalance, rather than choosing increasingly sophisticated approaches, and (ii) the ability to be trusted to be explained is a requirement of responsible working in both financial and blockchain settings. The research offers a pragmatic, interpretable framework on blockchain fraud detection and future directions, including sophisticated hybrid sampling, collective learning, as well as cross‐chain generalization to enhance fraud detection in distributed systems.