Related Experiment Video
Updated: Feb 14, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Using Large Language Models to Summarize Evidence in Biomedical Articles: Exploratory Comparison Between AI- and
Michelle Colder Carras1, Riaz Qureshi2, Kevin Naaman2
1Department of International Health, Johns Hopkins Bloomberg School of Public Health, 615 N Wolfe St, Baltimore, MD, 21205, United States, 1 410-955-3934.
AI-generated annotations using ChatGPT can provide a rapid overview of scientific literature, offering consistent quality and context but with potential inaccuracies and more errors than human annotations. Careful human verification is recommended for AI-generated annotated bibliographies.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Scientific Communication
Background:
- Annotated bibliographies are crucial for literature review but require significant human expertise and time.
- AI tools offer potential for efficient summarization, yet may introduce errors.
Purpose of the Study:
- To evaluate the feasibility of using ChatGPT for generating annotated bibliographies.
- To compare AI-generated annotations with human-written ones regarding accuracy, conciseness, and quality.
Main Methods:
- Two human annotators and three versions of ChatGPT (3.5, 4, 5) independently annotated 15 publications.
- Annotations were assessed for word count, readability (Flesch Reading Ease), capture of main points, presence of errors, and inclusion of quality/context.
- Statistical models were used to compare annotations generated by humans and AI.
Main Results:
- ChatGPT annotations were longer and less readable than human annotations.
- No significant difference was found in capturing main points between AI and human annotations.
- Human annotations had fewer errors, while AI annotations more consistently included quality and context, though sometimes inaccurately.
Conclusions:
- AI tools like ChatGPT can efficiently generate literature summaries, providing consistent quality and context.
- AI-generated annotations may contain more errors and inaccuracies, necessitating human verification.
- Further research into prompt engineering is needed to enhance AI chatbot performance for scientific literature summarization.
More Related Videos
Related Concept Videos
The Evidence for Evolution
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Genome Annotation and Assembly
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

