Related Experiment Video
Updated: Jun 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Capabilities of Generative AI Tools in Understanding Medical Papers: Qualitative Study
Seyma Handan Akyon1, Fatih Cagatay Akyon2,3, Ahmet Sefa Camyar4
1Golpazari Family Health Center, Bilecik, Turkey.
Large language models (LLMs) show varied performance in understanding medical research papers, with GPT-3.5-Turbo leading. This study evaluates LLM comprehension using the STROBE checklist for medical literature analysis.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Medical Informatics
Background:
- Medical literature comprehension is time-consuming for clinicians.
- Efficient tools are needed to help doctors process complex medical papers.
- Large language models (LLMs) offer potential solutions for medical information processing.
Purpose of the Study:
- To critically assess and compare the comprehension capabilities of various LLMs.
- To evaluate LLM accuracy and efficiency in understanding medical research papers.
- To utilize the STROBE checklist for a standardized evaluation framework.
Main Methods:
- Methodological research evaluating generative AI tools in medical paper comprehension.
- A benchmark pipeline processed 50 PubMed medical research papers.
- Six LLMs (GPT-3.5-Turbo, GPT-4, PaLM 2, Claude v1, Gemini Pro) were compared against expert benchmarks using 15 STROBE-derived questions.
Main Results:
- GPT-3.5-Turbo achieved the highest accuracy (66.9%), followed by GPT-4-1106 (65.6%) and PaLM 2 (62.1%).
- Statistically significant performance differences were observed among LLMs (P<.001).
- Newer LLM versions generally outperformed older ones, with notable performance variations across different paper sections.
Conclusions:
- This is the first study to use retrieval augmented generation to evaluate LLMs for medical paper comprehension.
- LLMs demonstrate potential to improve medical research efficiency and support evidence-based decision-making.
- Further research is required to address question format influence, potential biases, and the rapid evolution of LLM technology.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Non-equilibrium in the Cell
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Genomics
Patient-centered Care
Animal Mitochondrial Genetics
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...