Related Experiment Video
Updated: May 6, 2026

08:53
Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
5.1K
Large Language Models for Chatbot Health Advice Studies: A Systematic Review.
Bright Huo1, Amy Boyle2, Nana Marfo3
1Division of General Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
JAMA Network Open
|February 4, 2025
Summary
Reporting quality of studies on large language models (LLMs) in healthcare is inconsistent. This systematic review highlights the need for standardized reporting to ensure safe and effective AI integration in patient care.
Area of Science:
- Health Informatics
- Artificial Intelligence in Medicine
- Clinical Research Methodology
Background:
- Growing interest in integrating large language models (LLMs) into healthcare settings.
- Existing studies assessing LLM capabilities for health advice often lack consistent reporting quality.
- Need for standardized evaluation frameworks for AI-driven health tools.
Purpose of the Study:
- To systematically review reporting variability in peer-reviewed studies on generative AI chatbots for health advice.
- To identify gaps in reporting to inform the development of the Chatbot Assessment Reporting Tool (CHART).
Main Methods:
- Systematic literature search across MEDLINE, Embase, and Web of Science up to October 2023.
- Screening of 7752 articles by title/abstract and full text by two independent reviewers.
- Data extraction from 137 eligible studies evaluating clinical accuracy of AI chatbots for health advice.
Main Results:
- Included studies covered surgery, medicine, and primary care, focusing on treatment, diagnosis, and prevention.
- Most studies evaluated closed-source LLMs without specifying versions or key characteristics (e.g., temperature, token length).
- Reporting lacked details on prompt engineering, LLM querying dates, and objective performance metrics; ethical and safety implications were rarely addressed.
Conclusions:
- The reporting quality of studies on AI chatbots for health advice is heterogeneous.
- Findings underscore the need for standardized reporting guidelines, such as CHART, to ensure reliable evaluation.
- Addressing ethical, regulatory, and patient safety aspects is critical for clinical LLM integration.

