Related Experiment Video
Updated: Jan 17, 2026

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Does AI help humans make better decisions? A statistical evaluation framework for experimental and observational
Eli Ben-Michael1, D James Greiner2, Melody Huang3,4
1Department of Statistics & Data Science and Heinz College of Information & Systems Public Policy, Carnegie Mellon University, Pittsburgh, PA 15213.
Artificial intelligence (AI) does not improve judicial decision-making accuracy for cash bail. AI-alone systems, including large language models, performed worse than human judges.
Area of Science:
- Decision science
- Artificial intelligence
- Legal technology
Background:
- Data-driven algorithms, including artificial intelligence (AI), are increasingly used in society.
- Humans often retain final decision-making authority, especially in high-stakes situations.
- The effectiveness of AI in augmenting human decision-making remains a critical question.
Purpose of the Study:
- To introduce a methodological framework for empirically evaluating AI's impact on human decision-making.
- To compare the performance of human-alone, human-with-AI, and AI-alone decision systems.
- To determine if AI recommendations improve decision accuracy and when to follow them.
Main Methods:
- Developed a framework using standard classification metrics and potential outcomes.
- Employed a single-blinded, unconfounded treatment assignment for AI recommendations.
- Applied the methodology to a randomized controlled trial involving judges and a pretrial risk assessment instrument.
Main Results:
- AI-driven risk assessment recommendations did not enhance judges' accuracy in deciding on cash bail.
- AI-alone systems, including risk assessment scores and large language models, exhibited lower classification performance than human judges.
- The study provides insights into the optimal use of AI in human decision processes.
Conclusions:
- AI-generated recommendations did not improve judicial accuracy in pretrial risk assessment.
- Replacing human judges with AI systems, including large language models, resulted in diminished classification performance.
- The findings suggest caution in deploying AI for high-stakes legal decisions without rigorous empirical validation.
More Related Videos
07:42An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
Related Concept Videos
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
What is an Experiment?
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Experimental Designs