Related Experiment Video
Updated: May 6, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Automated labeling in medical data: A semi-supervised density-based approach for efficient diagnosis model
Lincy Meera Mathews1, Inaguri Muni Sai Haneesh1, S R Mani Sekhar1
1Department of Information Science and Engineering, M S Ramaiah Institute of Technology, Bangalore, India.
This study introduces a new semi-supervised learning method to automate medical data labeling, reducing costs and improving accuracy. The approach enhances diagnostic model development by efficiently utilizing unlabeled data.
Area of Science:
- Medical data analysis
- Machine learning in healthcare
- Computational biology
Background:
- Automated diagnosis and analysis models are crucial for healthcare practitioners due to the increasing volume of medical data.
- Manual data labeling for machine learning models is expensive, time-consuming, and prone to errors.
- Current methods struggle with the efficient utilization of large unlabeled medical datasets.
Purpose of the Study:
- To enhance semi-supervised learning performance by automating the data labeling process.
- To reduce the cost and complexity associated with developing automated medical diagnosis models.
- To improve the accuracy of diagnostic models using limited labeled data.
Main Methods:
- Developed a novel Semi-Supervised Density Based Clustering with a confidence region (SSDCAR) algorithm.
- Automated data labeling by identifying peak density samples and constructing clusters from unlabeled data.
- Analyzed sample distribution within clusters to identify high and low confidence regions for label propagation.
Main Results:
- The proposed SSDCAR algorithm demonstrated superior performance compared to existing methods on benchmark medical datasets.
- Achieved a significant accuracy increase of at least 2% across multiple health datasets.
- The algorithm proved to be scalable for larger datasets and memory-efficient.
Conclusions:
- SSDCAR offers a more accurate and efficient approach to semi-supervised learning in medical data analysis.
- The method effectively reduces the cost and manual effort required for data labeling.
- This technique provides a scalable and computationally efficient solution for developing automated diagnostic tools.
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Receiver Operating Characteristic Plot

