Related Experiment Video
Updated: May 9, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
461
Expert of Experts Verification and Alignment (EVAL) Framework for Large Language Models Safety in Gastroenterology.
Mauro Giuffrè1, Kisung You2, Ziteng Pang3
1Section of Digestive Diseases, Department of Medicine, Yale School of Medicine, New Haven, USA.
NPJ Digital Medicine
|May 3, 2025
Summary
We developed EVAL, a novel method to verify large language model (LLM) safety for medical queries. EVAL streamlines LLM output grading, improving accuracy for critical applications like upper gastrointestinal bleeding (UGIB).
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Clinical Decision Support
Background:
- Large language models (LLMs) offer potential for medical question answering but carry risks due to inaccurate outputs.
- Manual grading of LLM responses is impractical for clinical settings, hindering safe implementation.
- Ensuring LLM accuracy is crucial for high-stakes medical decision-making.
Purpose of the Study:
- To introduce EVAL (Expert-of-Experts Verification and Alignment), a system designed to streamline LLM output evaluation.
- To enhance the safety and accuracy of LLMs in medical applications, specifically for upper gastrointestinal bleeding (UGIB).
- To compare various LLM configurations and evaluation techniques.
Main Methods:
- Evaluated multiple LLMs (GPT, Claude, LLaMA, Mixtral) across 27 configurations, including zero-shot, retrieval-augmented generation, and supervised fine-tuning.
- Implemented EVAL using similarity-based ranking with Fine-Tuned ColBERT and a reward model trained on human-graded responses for rejection sampling.
- Assessed alignment with human performance using three separate datasets.
Main Results:
- Fine-Tuned ColBERT demonstrated the highest alignment with human performance (ρ = 0.81-0.91).
- The reward model accurately replicated human grading in 87.9% of cases.
- Rejection sampling using EVAL significantly improved LLM accuracy by 8.36% overall.
Conclusions:
- EVAL provides a scalable solution for assessing LLM accuracy in medical contexts.
- The developed system enhances LLM safety and reliability for clinical decision support.
- This approach is vital for the responsible deployment of AI in healthcare.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
2.5K
2.5K
Leaky Scanning
5.0K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.0K

