Related Experiment Video
Updated: Jan 14, 2026

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
9.9K
LegitPhish: A large-scale annotated dataset for URL-based phishing detection
Rachana S Potpelwar1, U V Kulkarni1, J M Waghmare1
1Shri Guru Gobind Singh Institute of Engineering and Technology, Vishnupuri, Nanded 431606, Maharashtra, India.
Data in Brief
|October 23, 2025
Summary
A new dataset, LegitPhish, offers 101,219 manually verified URLs for training machine learning models to detect phishing threats. This resource aids cybersecurity research and improves phishing detection accuracy.
Area of Science:
- Cybersecurity
- Machine Learning
- Data Science
Background:
- Phishing attacks pose a significant cybersecurity threat.
- Accurate and reliable datasets are crucial for developing effective phishing detection models.
- Existing datasets may lack manual verification or comprehensive features.
Purpose of the Study:
- Introduce LegitPhish, a novel, manually verified dataset of phishing and legitimate URLs.
- Facilitate research in machine learning-based phishing detection.
- Provide a benchmark for evaluating phishing detection systems.
Main Methods:
- Compiled a dataset of 101,219 labeled URLs (63,678 phishing, 37,540 legitimate).
- Manually verified all URLs for accuracy.
- Annotated URLs with 17 structural and lexical features.
- Sourced phishing URLs from threat intelligence feeds and legitimate URLs from high-authority domains.
Main Results:
- LegitPhish contains 101,219 manually verified URLs.
- The dataset includes 17 diverse features for each URL.
- Provides a robust resource for machine learning model training and evaluation.
Conclusions:
- LegitPhish is a valuable, publicly available resource for cybersecurity research.
- The dataset supports reproducible research and benchmarking in phishing detection.
- Enhances the development of more reliable web security applications.
