Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Disparities in AI-Based Prior Authorization for Head and Neck Reconstruction: A Large Language Model Analysis
Shannon S Wu1, Mugil V Shanmugam2, Yu-Jin Lee1
1Department of Otolaryngology-Head and Neck Surgery Stanford University School of Medicine Palo Alto California USA.
Summary
Large language models (LLMs) show bias in insurance authorization for cancer reconstruction, favoring certain demographics. More data and safeguards are needed for equitable AI in healthcare.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Health Equity Research
Background:
- Large language models (LLMs) are increasingly used in healthcare for administrative tasks.
- The potential for bias in AI-driven insurance authorization decisions is not well understood.
- This study investigates bias in LLM-based prior authorization for otolaryngology, focusing on oral cavity squamous cell carcinoma (SCC).
Purpose of the Study:
- To assess bias in Large Language Model (LLM)-driven prior authorization for head and neck cancer reconstruction.
- To simulate insurance coverage decisions using GPT-4o for oral cavity squamous cell carcinoma (SCC) treatment.
- To identify demographic disparities in LLM recommendations for surgical reconstruction.
Main Methods:
- Utilized OpenAI's GPT-4o to simulate 19,900 insurance authorization decisions for T2N2 oral cavity SCC.
- Compared two reconstruction methods: radial forearm free flap (RFFF) and split-thickness skin graft (STSG).
- Systematically varied patient profiles by age, sex, race/ethnicity, income, socioeconomic status (SES), and substance use.
Main Results:
- Significant disparities observed: RFFF was favored for younger, Asian, or white patients from high-income/SES backgrounds (p < 0.0001).
- Older, Black, and Hispanic patients, and those from lower-income areas or with substance use history, received fewer RFFF approvals (p < 0.0001).
- Including tumor-specific data in prompts reduced sociodemographic bias, favoring RFFF across groups.
Conclusions:
- LLM outputs demonstrated significant bias based on patient demographics when clinical information was limited.
- Detailed and pertinent clinical data are crucial for LLM inputs to mitigate bias in decision-making.
- Increased regulatory oversight, bias recognition, and safeguards are essential for equitable AI in insurance authorization.