Related Experiment Video
Updated: Jan 18, 2026

Failure of Cleaning Verification in Pharmaceutical Industry Due to Uncleanliness of Stainless Steel Surface
Published on: August 11, 2017
A two-stage active cleaning strategy for long-tail label noise.
Xiao Lin1, Zeyu Rong2, Yan Li3
1The College of Information, Mechanical and Electrical Engineering, Shanghai Normal University, Shanghai, China; Shanghai Intelligent Education Big Data Engineering Technology Research Center, Shanghai Normal University, Shanghai, China; Lab for Educational Big Data and Policymaking, Ministry of Education, Shanghai Normal University, Shanghai, China; Shanghai Online Education Research Base for Primary and Secondary Schools, Shanghai, China.
We introduce a two-stage active label cleaning strategy to efficiently handle long-tailed data with label noise. This method enhances feature representations and uses active learning to minimize re-annotation costs, improving classification performance.
Area of Science:
- Machine Learning
- Computer Vision
- Data Science
Background:
- Long-tailed data presents challenges in real-world applications due to imbalanced classes and label noise.
- Existing methods for noisy long-tailed data often require high computational and manual annotation efforts.
Purpose of the Study:
- To propose a novel, cost-effective two-stage active label cleaning strategy for long-tailed datasets with label noise.
- To enhance feature representation quality and efficiently identify and correct mislabeled samples.
Main Methods:
- Stage 1: Balanced Class-Centered Contrastive Learning (BCCL) to improve feature representations and detect potential label noise.
- Stage 2: Uncertainty-based active learning for focused re-labeling of high-uncertainty samples.
- Iterative re-labeling process to refine classification performance and optimize annotation resources.
Main Results:
- The proposed method demonstrates robustness across various noise ratios and imbalance levels.
- Significantly outperforms state-of-the-art methods, achieving a 5.17% relative improvement on CIFAR10-LT.
- Achieves 38.37% accuracy on the Red Mini-ImageNet dataset, outperforming existing baselines in high-noise scenarios.
Conclusions:
- The two-stage active label cleaning strategy effectively improves classification performance on noisy, long-tailed data.
- Minimizes annotation workload and optimizes resource utilization through efficient sample selection for re-labeling.
- Offers a superior solution for handling challenging real-world datasets with label noise and class imbalance.

