Related Experiment Video
Updated: Feb 22, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Meaningless comparisons lead to false optimism in medical machine learning
Orianna DeMasi1, Konrad Kording2,3, Benjamin Recht1
1Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, California, United States of America.
Many digital health algorithms overestimate their effectiveness by using weak comparisons. This study introduces "user lift" to provide a more accurate evaluation of personalized medical monitoring tools.
Area of Science:
- Digital health
- Medical informatics
- Mental wellbeing monitoring
Background:
- Algorithms analyzing big datasets from personal devices are a growing trend in medicine.
- Current evaluation methods for these algorithms often use weak baselines, leading to inflated performance metrics.
- This is particularly concerning in mental wellbeing monitoring where patient-specific baselines are crucial.
Purpose of the Study:
- To meta-analyze the literature on algorithm evaluation for mental wellbeing monitoring.
- To identify and quantify the prevalence of inadequate baseline comparisons.
- To propose a novel metric, "user lift," to address systematic errors in algorithm evaluation.
Main Methods:
- Meta-analysis of studies evaluating algorithms for mental wellbeing monitoring.
- Quantitative assessment of baseline comparison methodologies used in the literature.
- Development and proposal of the "user lift" metric.
Main Results:
- Approximately 77% of the literature employs inadequate comparisons that ignore patient baseline states.
- Predicting a patient's average state can explain a large portion of variance, rendering simple algorithms seemingly effective but practically useless.
- Inappropriate baseline comparisons lead to significant overestimation of algorithm performance and "baseless optimism" in the field.
Conclusions:
- Current evaluation practices for digital health algorithms in mental wellbeing monitoring are flawed.
- The proposed "user lift" metric offers a more robust method for assessing personalized medical monitoring.
- Accurate algorithm evaluation is essential to avoid misleading conclusions and ensure the clinical utility of digital health tools.
Related Concept Videos
Regression Toward the Mean
Unrealistic Optimism Bias
Errors occurring during blood pressure monitoring
Several factors...
Improving Translational Accuracy
Improving Translational Accuracy
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
