Related Experiment Video
Updated: Oct 5, 2025

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
Peeking into a black box, the fairness and generalizability of a MIMIC-III benchmarking model
Eliane Röösli1,2, Selen Bozkurt2, Tina Hernandez-Boussard3,4
1School of Life Sciences, Swiss Federal Institute of Technology (EPFL), Lausanne, Switzerland.
Abstract:
As artificial intelligence (AI) makes continuous progress to improve quality of care for some patients by leveraging ever increasing amounts of digital health data, others are left behind. Empirical evaluation studies are required to keep biased AI models from reinforcing systemic health disparities faced by minority populations through dangerous feedback loops. The aim of this study is to raise broad awareness of the pervasive challenges around bias and fairness in risk prediction models. We performed a case study on a MIMIC-trained benchmarking model using a broadly applicable fairness and generalizability assessment framework. While open-science benchmarks are crucial to overcome many study limitations today, this case study revealed a strong class imbalance problem as well as fairness concerns for Black and publicly insured ICU patients. Therefore, we advocate for the widespread use of comprehensive fairness and performance assessment frameworks to effectively monitor and validate benchmark pipelines built on open data resources.
Related Concept Videos
Modeling and Similitude
Self-Evaluation Maintenance Model
Regression Toward the Mean
Stereotype Content Model
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Strategies of Self-Presentation III: Self-Monitoring

