Related Experiment Video
Updated: Mar 29, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Applications of Large Language Models in Medical Research: From Systematic Reviews to Clinical Studies
Eun Jeong Gong1,2,3, Chang Seok Bang1,2,3, Yong Seok Shin1
1Department of Internal Medicine, Hallym University College of Medicine, Chuncheon 24252, Republic of Korea.
Bioengineering (Basel, Switzerland)
|March 28, 2026
Summary
Large Language Models (LLMs) show promise in medical research, aiding systematic reviews and clinical tasks. However, their instability and potential for bias necessitate human oversight for reliable scientific advancement.
Area of Science:
- Medical Research
- Artificial Intelligence
- Scientific Writing
Background:
- Large Language Models (LLMs) are increasingly integrated into medical research.
- Their impact spans systematic reviews, scientific writing, and clinical research workflows.
Purpose of the Study:
- To synthesize current evidence on the applications of LLMs in medical research.
- To evaluate the efficacy and limitations of LLMs in various research contexts.
Main Methods:
- A narrative review of literature published between 2023-2025.
- Searches conducted across major scientific databases (PubMed, Scopus, Web of Science, arXiv, medRxiv, Google Scholar).
- Inclusion of studies with empirical findings or methodological evaluations of LLM applications.
Main Results:
- LLMs demonstrate high accuracy (80-94%) in systematic review data extraction but show limited agreement (κ = 0.16-0.43) in risk-of-bias assessment.
- Scientific writing applications exhibit significant hallucination rates (47-55%) and demographic bias (>90%), requiring stringent verification.
- LLMs assist in clinical research tasks like coding and protocol development, but human validation is essential.
Conclusions:
- LLMs are powerful yet unstable tools in medical research, demanding continuous human verification.
- Maintaining a human-in-the-loop approach is crucial to balance AI efficiency with critical thinking and prevent cognitive offloading.
Keywords:
ChatGPTGPT-4artificial intelligenceevidence synthesislarge language modelsmedical researchprompt engineeringsystematic review
