Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Disparities in AI-Based Prior Authorization for Head and Neck Reconstruction: A Large Language Model Analysis
Shannon S Wu1, Mugil V Shanmugam2, Yu-Jin Lee1
1Department of Otolaryngology-Head and Neck Surgery Stanford University School of Medicine Palo Alto California USA.
Summary
Large language models (LLMs) used for insurance authorization show bias, favoring certain demographics for cancer reconstruction. Addressing AI bias and improving data inputs are crucial for equitable healthcare.
Area of Science:
- Artificial Intelligence in Healthcare
- Health Equity and Bias
- Otolaryngology and Oncology
Background:
- Large language models (LLMs) are increasingly used in healthcare for administrative tasks, including insurance authorization.
- The potential for AI-driven systems to perpetuate or introduce bias in healthcare decision-making, particularly for insurance coverage, is a growing concern.
- Prior authorization in otolaryngology, especially for complex cases like head and neck cancer reconstruction, is a critical area where AI implementation requires scrutiny.
Purpose of the Study:
- To examine potential bias in large language model (LLM)-driven prior authorization processes within otolaryngology.
- To investigate LLM performance in simulating insurance coverage decisions for oral cavity squamous cell carcinoma (SCC) reconstruction.
- To assess how patient demographic variables influence LLM-based coverage recommendations.
Main Methods:
- Utilized OpenAI's GPT-4o to simulate 19,900 insurance coverage decisions for head and neck cancer reconstruction.
- Developed a standardized clinical scenario for T2N2 oral cavity SCC requiring surgical resection and reconstruction, comparing radial forearm free flap (RFFF) vs. split-thickness skin graft (STSG).
- Systematically varied patient profiles by age, sex, race/ethnicity, income level, socioeconomic status (SES), and substance use history to test for bias.
Main Results:
- Significant disparities in LLM approval decisions were observed, with RFFF more frequently approved for younger, Asian, or white patients from high-income/SES backgrounds (p < 0.0001).
- Older, Black, and Hispanic patients, and those from lower-income areas or with substance use histories, were less likely to receive RFFF authorization (p < 0.0001).
- Sensitivity analysis indicated that including tumor-specific information significantly skewed recommendations towards RFFF across all sociodemographic groups.
Conclusions:
- LLM outputs demonstrated significant bias based on patient demographics when clinical information was limited, impacting oral cavity cancer reconstruction recommendations.
- To mitigate bias in AI-driven clinical decision-making, it is essential to incorporate comprehensive and pertinent patient information.
- Increased regulatory governance, rigorous safeguards, and awareness of AI biases are necessary as insurers adopt AI for prior authorization to ensure equitable healthcare.