Related Experiment Video
Updated: Jun 24, 2025

Author Spotlight: Developing a Point-of-Care Hemoglobin Estimation Method for Anemia Management
Published on: January 19, 2024
Evaluating accuracy and fairness of clinical decision support algorithms when health care resources are limited
Esther L Meerwijk1, Duncan C McElfresh2, Susana Martins2
1Program Evaluation and Resource Center, Office of Mental Health and Suicide Prevention, Department of Veterans Affairs, Menlo Park, CA, USA; VA Health Systems Research, Center for Innovation to Implementation (Ci2i), VA Palo Alto Health Care System, Menlo Park, CA, USA.
Objective:
Guidance on how to evaluate accuracy and algorithmic fairness across subgroups is missing for clinical models that flag patients for an intervention but when health care resources to administer that intervention are limited. We aimed to propose a framework of metrics that would fit this specific use case.
Methods:
We evaluated the following metrics and applied them to a Veterans Health Administration clinical model that flags patients for intervention who are at risk of overdose or a suicidal event among outpatients who were prescribed opioids (N = 405,817): Receiver - Operating Characteristic and area under the curve, precision - recall curve, calibration - reliability curve, false positive rate, false negative rate, and false omission rate. In addition, we developed a new approach to visualize false positives and false negatives that we named 'per true positive bars.' We demonstrate the utility of these metrics to our use case for three cohorts of patients at the highest risk (top 0.5 %, 1.0 %, and 5.0 %) by evaluating algorithmic fairness across the following age groups: <=30, 31-50, 51-65, and >65 years old.
Results:
Metrics that allowed us to assess group differences more clearly were the false positive rate, false negative rate, false omission rate, and the new 'per true positive bars'. Metrics with limited utility to our use case were the Receiver - Operating Characteristic and area under the curve, the calibration - reliability curve, and the precision - recall curve.
Conclusion:
There is no "one size fits all" approach to model performance monitoring and bias analysis. Our work informs future researchers and clinicians who seek to evaluate accuracy and fairness of predictive models that identify patients to intervene on in the context of limited health care resources. In terms of ease of interpretation and utility for our use case, the new 'per true positive bars' may be the most intuitive to a range of stakeholders and facilitates choosing a threshold that allows weighing false positives against false negatives, which is especially important when predicting severe adverse events.
More Related Videos
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Ethical Issues
Ethical Concerns in Healthcare:
Ethical Dilemmas II
Methods Of Healthcare Delivery System
Managed Care System:
The managed care system is designed to control the cost while maintaining the quality of care. The patient's care from admission to discharge is planned by the primary care provider or the case manager, also known as the gatekeeper. In a managed care system, the number of care providers is...

