Related Experiment Video
Updated: Apr 19, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Learning to improve medical decision making from imbalanced data without a priori cost
Xiang Wan1, Jiming Liu2, William K Cheung3
1Department of Computer Science and Institute of Computational and Theoretical Studies, Hong Kong Baptist University, Kowloon Tong, Hong Kong. xwan@comp.hkbu.edu.hk.
RankCost effectively classifies imbalanced medical data without needing prior cost information. This novel approach consistently performs well across various datasets, offering a reliable solution for medical decision-making challenges.
Area of Science:
- Medical Data Analysis
- Machine Learning
- Bioinformatics
Background:
- Medical datasets often exhibit imbalanced class distributions, with a rare abnormal group and a common normal group.
- Misclassifying rare abnormal cases as normal incurs significant costs, a key challenge in medical diagnosis.
- Traditional classification methods struggle with skewed data, often relying on difficult-to-estimate a priori costs.
Purpose of the Study:
- To introduce a novel machine learning method, RankCost, for classifying imbalanced medical data.
- To address the challenge of imbalanced classification without requiring prior cost information.
- To develop a robust method for medical decision-making applications.
Main Methods:
- RankCost transforms imbalanced classification into a partial ranking problem using a scoring function.
- The scoring function aims to maximize the separation between minority and majority classes.
- A non-parametric boosting algorithm is employed to learn the scoring function.
Main Results:
- RankCost was compared against several existing methods on diverse medical datasets.
- The proposed method demonstrated consistent and comparable performance across datasets with varying sizes and imbalance ratios.
- Unlike other methods, RankCost's performance was not dependent on specific a priori cost values.
Conclusions:
- Learning effective classification models from imbalanced medical data remains a significant challenge.
- RankCost offers a novel solution by avoiding the need for a priori cost estimation.
- Experimental results confirm RankCost's efficacy and potential utility in real-world medical decision-making.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Kaplan-Meier Approach
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...