Related Experiment Video
Updated: Sep 20, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Decisions are not all equal-Introducing a utility metric based on case-wise raters' perceptions
Andrea Campagner1, Federico Sternini2, Federico Cabitza3
1Dipartimento di Informatica, Sistemistica e Comunicazione, Università di Milano-Bicocca, Milano, Italy.
A new weighted Utility (wU) metric enhances AI decision support system (AI-DSS) evaluation by incorporating human perception of case relevance and annotation hesitation. This human-centered approach improves AI model performance, especially for complex cases.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Machine Learning
Background:
- Evaluating AI-based decision support systems (AI-DSS) is crucial but current metrics often neglect contextual information.
- Existing metrics may not fully capture the nuances of human-AI interaction and decision-making processes.
Purpose of the Study:
- Introduce a novel utility metric, weighted Utility (wU), for AI-DSS evaluation.
- Demonstrate the application of wU for both AI model evaluation and optimization.
- Highlight the importance of human-centered evaluation in critical domains.
Main Methods:
- Developed the weighted Utility (wU) metric based on rater perceptions of annotation hesitation and training case relevance.
- Compared wU with existing metrics like Net Benefit and error-based metrics.
- Applied wU in three realistic case studies for model evaluation and optimization.
Main Results:
- The wU metric generalizes existing evaluation metrics.
- wU provides a more flexible tool for AI model evaluation.
- Optimization using wU significantly improved model performance (AUC 0.862 vs 0.895, p<0.05), particularly for complex cases (AUC 0.85 vs 0.92, p<0.05).
Conclusions:
- Utility should be a primary concern in evaluating and optimizing machine learning models in critical fields like medicine.
- A human-centered approach is vital for assessing AI's impact on human decision-making.
- Information gathered during ground-truthing can enhance AI model evaluation.
More Related Videos
06:43The Crossmodal Congruency Task as a Means to Obtain an Objective Behavioral Measure in the Rubber Hand Illusion Paradigm
Published on: July 26, 2013
07:05Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Related Concept Videos
Ratio Level of Measurement
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Odds Ratio
Review and Preview
Percentiles are a type of fractile that partition data into...
Ranks