Related Experiment Video
Updated: Sep 16, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Framework for bias evaluation in large language models in healthcare settings
Tara Templin1,2,3, Sophia Fort4, Prasanna Padmanabham5
1Department of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA. ttemplin@unc.edu.
A new audit framework addresses the lack of standardized evaluation for AI clinical decision models, ensuring accuracy and fairness. This approach focuses on model outputs for responsible AI adoption in healthcare.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Healthcare Technology Evaluation
Background:
- Large language models (LLMs) offer potential for AI-assisted clinical decisions but lack standardized evaluation methods.
- Current adoption is hindered by concerns regarding model accuracy, bias, and a lack of transparent auditing processes.
- Ensuring the reliability and fairness of AI tools in healthcare is paramount for patient safety and effective clinical practice.
Purpose of the Study:
- To introduce a novel, standardized five-step audit framework for evaluating LLMs used in clinical decision support.
- To guide healthcare practitioners in assessing AI model accuracy and identifying potential biases.
- To promote the responsible development and deployment of AI in clinical settings.
Main Methods:
- Development of a five-step framework encompassing stakeholder engagement, population-specific model calibration, and scenario-based testing.
- Creation of open-access tools to facilitate stakeholder involvement in the audit process.
- Demonstration of the framework's application through a practical audit example.
Main Results:
- The proposed framework provides a structured approach to auditing AI clinical decision models.
- It emphasizes testing model outputs against clinically relevant scenarios to ensure performance and identify bias.
- Open-access tools and an example audit are provided to aid practitioners.
Conclusions:
- A standardized audit framework is critical for the safe and effective adoption of LLMs in clinical decision-making.
- Focusing audits on model outputs rather than internal parameters promotes responsible AI use.
- This framework supports regulatory efforts and builds trust in AI healthcare applications.
More Related Videos
Related Concept Videos
Bias in Epidemiological Studies
Improving Translational Accuracy
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Nursing Evaluation
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...

