Related Experiment Video
Updated: Aug 22, 2025

Automated Detection and Analysis of Exocytosis
Published on: September 11, 2021
AndroMalPack: enhancing the ML-based malware classification by detection and removal of repacked apps for Android
Husnain Rafiq1, Nauman Aslam2, Muhammad Aleem3
1Department of Computer and Information Sciences, Northumbria University, Newcastle upon Tyne, UK. husnain.rafiq@northumbria.ac.uk.
This study addresses the problem of duplicate malicious software in Android datasets. By identifying and removing repacked apps, the researchers developed a more efficient detection tool that maintains high accuracy while reducing training data requirements.
Area of Science:
- Cybersecurity research within Android malware classification
- Machine learning applications in mobile device security
Background:
Mobile device security faces persistent challenges from malicious software proliferation. Prior research has shown that public repositories often contain redundant entries. This gap motivated an investigation into the prevalence of cloned applications. It was already known that developers frequently modify existing malicious code to evade detection. That uncertainty drove the need for quantifying how many samples represent unique threats. No prior work had resolved the exact proportion of duplicates across major benchmark collections. Researchers previously lacked a standardized approach to filter these repetitive files effectively. This study provides a necessary foundation for improving automated threat identification systems.
Purpose Of The Study:
The aim of this research is to enhance machine learning-based detection by filtering redundant malicious applications. The authors address the issue of inflated performance metrics caused by dataset cloning. This problem complicates the development of robust security tools for mobile platforms. The researchers seek to quantify the exact extent of this redundancy in common benchmarks. They intend to demonstrate that cleaner data leads to more reliable threat identification. The team also strives to develop an optimized detector that functions effectively with fewer samples. This effort focuses on improving the efficiency of training processes for security software. The study ultimately seeks to provide a standardized approach for future malware analysis.
Main Methods:
The review approach involves a systematic evaluation of three prominent Android security benchmarks. Researchers processed 5560 samples from the Drebin collection for initial assessment. They expanded the scope to include 24,533 entries from the AMD repository. The team also examined 695,470 files sourced from the AndroZoo archive. Investigators applied package name comparisons to detect structural similarities between distinct files. This strategy enabled the isolation of redundant malicious software versions. The authors then constructed a specialized detector using optimized machine learning architectures. They finalized the process by validating the tool against diverse, filtered testing environments.
Main Results:
Key Findings From the Literature indicate that a large portion of benchmark data consists of simple clones. The analysis shows that 52.3% of the Drebin dataset contains repacked software. Researchers observed that 29.8% of the AMD collection is composed of duplicate entries. The study reveals that 42.3% of the AndroZoo repository consists of identical malicious variants. The proposed detector achieves a detection accuracy reaching 98.2% on cleaned datasets. The system maintains exceptionally low false-positive rates during performance testing. These results demonstrate that smaller, high-quality training sets outperform larger, redundant ones. The evidence confirms that nature-inspired optimization effectively improves classification outcomes for mobile threats.
Conclusions:
The authors demonstrate that removing redundant samples improves model training efficiency. Synthesis and Implications suggest that dataset quality significantly impacts classification performance. The researchers propose that their new detector maintains high precision despite using smaller training sets. This work confirms that nature-inspired optimization techniques enhance detection capabilities for mobile threats. The findings indicate that current benchmarks likely overestimate the diversity of available malicious samples. The team suggests that future studies should prioritize the use of cleaned, clone-free data. This analysis provides a framework for researchers to standardize their evaluation protocols. The published list of identified clones serves as a resource for the broader security community.
Frequently Asked Questions
The researchers propose that AndroMalPack utilizes nature-inspired algorithms to optimize classification. This approach allows the system to achieve a 98.2% detection accuracy while maintaining low false-positive rates, even when trained on a significantly reduced set of unique malicious applications.
The authors utilize package names-based similarity to identify clones. This method compares the unique identifiers of applications across the Drebin, AMD, and AndroZoo datasets to quantify the extent of repacked software present in these widely used security benchmarks.
The researchers emphasize that cleaning datasets is necessary because over 50% of some collections consist of clones. By removing these duplicates, the model avoids overfitting to repetitive patterns, which is essential for accurately identifying novel threats in the wild.
The authors use three large-scale datasets, specifically Drebin, AMD, and AndroZoo, to validate their findings. These collections serve as the data type for quantifying the prevalence of clones and training the new detector.
The study measures the prevalence of repacked malware, finding 52.3% in Drebin, 29.8% in AMD, and 42.3% in AndroZoo. This phenomenon highlights a significant redundancy issue in current cybersecurity research benchmarks.
The authors claim that their work fosters future research by providing a public dataset of cloned apps. They suggest this contribution will help other investigators improve the reliability of their own malware detection models.
Related Concept Videos
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Stereotype Content Model
Methods of Classification and Identification
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Viral Recombination

