Related Experiment Video
Updated: Jan 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated Literature Screening for Hepatocellular Carcinoma Treatment Through Integration of 3 Large Language Models:
Chen Pan1, Wei Lu1, Bingliang Chen1
1Department of Hepatobiliary and Vascular Surgery, First Affiliated Hospital of Chengdu Medical College, Chengdu, China.
This study introduces an automated literature screening system using large language models (LLMs) to speed up evidence synthesis for hepatocellular carcinoma (HCC) treatment guidelines. The LLM framework accelerates guideline updates by reducing screening time while maintaining rigor.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Oncology Research
Background:
- Primary liver cancer, particularly hepatocellular carcinoma (HCC), presents diagnostic and therapeutic challenges.
- Timely updates to clinical guidelines are crucial but hindered by the labor-intensive nature of systematic reviews.
- Evidence synthesis for HCC treatment requires efficient and accurate literature screening methods.
Purpose of the Study:
- To develop and evaluate an automated literature screening workflow using large language models (LLMs) for accelerating evidence synthesis in HCC treatment guidelines.
- To simulate collaborative decision-making for study inclusion/exclusion using a tripartite LLM framework.
- To assess the performance and efficiency of the LLM-driven approach compared to traditional methods.
Main Methods:
- A tripartite LLM framework was developed, integrating three distinct models (Doubao-1.5-pro-32k, Deepseek-v3, DeepSeek-R1-Distill-Qwen-7B).
- The framework was evaluated on 9 reconstructed datasets from published HCC meta-analyses.
- Performance metrics included accuracy, agreement (κ, prevalence-adjusted bias-adjusted κ), recall, precision, F1-scores, processing time, and cost.
Main Results:
- The LLM framework achieved a weighted accuracy of 0.96 and substantial agreement (prevalence-adjusted bias-adjusted κ=0.91).
- High weighted recall (0.90) was observed, alongside modest weighted precision (0.15) and F1-scores (0.22).
- Computational efficiency showed variation, with processing times ranging from 248-5850 seconds and costs from $0.14-$3.68 per dataset.
Conclusions:
- The LLM-driven approach shows promise for accelerating evidence synthesis in HCC care, reducing screening time while maintaining methodological rigor.
- Limitations include clinical context sensitivity and error propagation, suggesting the need for reinforcement learning and domain-specific fine-tuning.
- LLM agent architectures with reinforcement learning offer a practical path for streamlining guideline updates, requiring further optimization for reliability in complex clinical settings.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025