Related Experiment Video
Updated: Jun 20, 2026

05:33
Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Hepatology e-consult responses generated by artificial intelligence demonstrate accuracy but require human oversight
Holly K T Huang1, Debra W Yen2, Michelle Y Li1
1Division of Gastroenterology and Hepatology, Department of Medicine, University of California San Francisco, San Francisco, California, USA.
Hepatology Communications
|June 19, 2026
Summary
A customized large language model (LLM) for liver diseases, LiVersa, shows promise in drafting electronic consultations (e-consults). While effective, human oversight is crucial, and LLM-as-a-judge offers conservative quality assurance.
Area of Science:
- Hepatology
- Artificial Intelligence
- Medical Informatics
Background:
- Electronic consultations (e-consults) enhance specialist access but increase provider workload.
- A novel customized large language model (LLM), LiVersa, was developed for liver disease e-consults.
- The study evaluates LiVersa's efficacy in drafting responses and compares human and machine reviewer assessments.
Purpose of the Study:
- To assess the performance of LiVersa in generating hepatology e-consult responses.
- To evaluate the equivalence between human expert reviewers and an LLM-as-a-judge system.
- To determine the accuracy, safety, and usability of LLM-generated e-consult drafts.
Main Methods:
- LiVersa-generated responses for 61 hepatology e-consults from UCSF (Jan-Mar 2025) were analyzed.
- Three hepatologists and an "LLM-as-a-judge" system evaluated drafts using a 12-item rubric.
- Equivalence testing between human and LLM reviewers was performed using two one-sided tests (TOST).
Main Results:
- LiVersa drafts matched original responses in word count and verbosity.
- Human reviewers found 72% of drafts reasonable starting points and 83% provided appropriate recommendations.
- LLM-as-a-judge was more stringent, identifying more potentially harmful drafts (67% vs. 20%) and fewer clinically equivalent responses (27% vs. 48%) than human reviewers.
Conclusions:
- Customized LLMs like LiVersa demonstrate potential for streamlining e-consult drafting.
- Human oversight remains essential for ensuring the safety and accuracy of LLM-generated responses.
- LLM-as-a-judge serves as a conservative tool for quality assurance, particularly during LLM updates.