Related Experiment Video
Updated: May 21, 2025

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.3K
Explainable TabNet ensemble model for identification of obfuscated URLs with features selection to ensure secure web
Mehwish Naseer1, Farhan Ullah2, Saqib Saeed3
1Computer and Software Engineering Department, College of Electrical and Mechanical Engineering, National University of Sciences and Technology (NUST), Islamabad, 44080, Pakistan.
Scientific Reports
|March 20, 2025
Summary
This study introduces a robust TabNet ensemble model to accurately identify malicious URLs, offering enhanced cybersecurity against threats like malware and phishing. The model achieved high performance metrics, demonstrating its effectiveness in classifying harmful web addresses.
Area of Science:
- Cybersecurity
- Machine Learning
- Artificial Intelligence
Background:
- Obfuscated and malicious Uniform Resource Locators (URLs) pose significant cybersecurity risks, including malware distribution, phishing, and scams.
- Effective identification of malicious URLs is a critical challenge in safeguarding systems and users.
Purpose of the Study:
- To propose a Robust unified TabNet ensemble model for the accurate identification and classification of malicious URLs.
- To leverage feature importance for enhanced classification performance.
Main Methods:
- Utilized a fine-tuned attention-based deep neural network (TabNet) for URL feature extraction.
- Developed a Machine Learning (ML) ensemble model using customized data with the most important features.
- Employed statistical validation including Kappa value and 10-fold cross-validation.
- Applied Local Interpretable Model-agnostic Explanations (LIME) for model interpretability.
Main Results:
- Achieved high performance with 97.8% accuracy, 0.978 precision, 0.976 recall, and 0.978 F1-score for classifying five URL classes.
- Statistical validation showed a Kappa value of 0.968.
- 10-fold cross-validation yielded a mean accuracy of 97.27% with a confidence interval of 0.004.
Conclusions:
- The proposed TabNet ensemble model demonstrates superior efficacy in malicious URL identification compared to state-of-the-art methods.
- The model's interpretability through LIME validates the contribution of key features in classification.
- The findings support the model's potential for practical application in cybersecurity defense.

