Related Experiment Video
Updated: Jul 5, 2025

08:05
Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
10.7K
Methodology for the Detection of Contaminated Training Datasets for Machine Learning-Based Network
Joaquín Gaspar Medina-Arco1, Roberto Magán-Carrión1, Rafael Alejandro Rodríguez-Gómez1
1Network Engineering & Security Group (NESG), University of Granada, 18012 Granada, Spain.
Sensors (Basel, Switzerland)
|January 23, 2024
Summary
This study introduces a new method to improve network intrusion detection systems (NIDS) by identifying and correcting mislabelled data in training sets, enhancing anomaly detection accuracy against cyber threats.
Area of Science:
- Cybersecurity
- Machine Learning
- Network Security
Background:
- Network Intrusion-Detection Systems (NIDS) are crucial for detecting cyber-attacks.
- Anomaly-based NIDS rely on machine learning models trained on labelled datasets.
- Mislabeled data in training sets can significantly degrade NIDS performance.
Purpose of the Study:
- To address the challenge of mislabelled network traffic datasets in anomaly-based NIDS.
- To develop a methodology for analyzing dataset quality and identifying hidden anomalies.
- To optimize NIDS performance by selecting ideal training subsets, even with mislabelled data.
Main Methods:
- A novel two-step methodology is proposed to analyze network traffic dataset quality.
- The method identifies hidden or unidentified anomalies within existing datasets.
- It involves an incremental selection of data subsets to train anomaly detection models.
Main Results:
- The proposed methodology successfully identified hidden botnet attacks in the contaminated UGR'16 dataset.
- Experiments demonstrated the feasibility of revealing mislabelled data, including normal traffic labelled as anomalous.
- The approach improved the performance of a state-of-the-art NIDS (Kitsune) on the compromised dataset.
Conclusions:
- The developed methodology effectively enhances the robustness of anomaly-based NIDS against dataset mislabelling.
- It provides a reliable way to improve the accuracy of intrusion detection systems.
- This approach is crucial for maintaining effective network security in the face of evolving cyber threats.
Related Concept Videos
Survival Tree
86
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
86
Steps in Outbreak Investigation
131
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
131

