LegitPhish: A large-scale annotated dataset for URL-based phishing detection

Rachana S. Potpelwar, Uday V. Kulkarni, Jaishri M. Waghmare · Data in Brief · 2025

Phishing attacks are a major cybersecurity threat, requiring timely detection using reliable datasets. We present LegitPhish, a novel and manually verified dataset of phishing and legitimate URLs, designed to facilitate research in machine learning-based phishing detection. LegitPhish, a publicly available dataset of 101,219 labelled URLs, including 63,678 phishing and 37,540 legitimate entries. All URLs are manually verified and annotated with 17 structural and lexical features such as URL length, token count, entropy, subdomain usage, and TLD characteristics. Phishing URLs were collected from threat intelligence feeds (URL Haus, Phish Tank) [1,2] and verified for accuracy, while legitimate URLs were sourced using Google Search API and curated from high-authority domains like Wikipedia. LegitPhish enables reproducible research and benchmarking for phishing detection models and web security applications.

Read the paper · More papers on PaperTik