Privacy-Preserving Phishing Detection in HTML Code Using Split Learning
Jung In Kim, Yushin Kim, Sejong Lee, Sunghyun Cho · 2024
Split learning is a distributed learning technique that enables multiple data owners to collaboratively train deep learning models without sharing their raw data, thereby pre-serving privacy and reducing computational burdens. This study investigates the application of split learning in the domain of phishing detection using HyperText Mark-up Language (HTML) code analysis. By integrating large language model (LLM) within a split learning framework, we aim to enhance the detection of phishing attempts while maintaining data privacy and optimizing computational resources. We conducted extensive experiments comparing the performance of the LLM split learning model with traditional centralized models, assessing scenarios with biased client sampling, varying client numbers, and pre-trained server models. The results indicate that the split learning model achieves comparable accuracy to centralized models, demonstrating its robustness and efficiency. Our findings underscore the potential of split learning in developing privacy-preserving and computationally efficient anti-phishing systems.