Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Methods of Documentation III: PIE01:21

Methods of Documentation III: PIE

1.5K
Problem-intervention-evaluation (PIE) is a systematic approach to documentation used in healthcare settings for clinical decision-making and patient care planning. It is a structured approach to organizing patient data based on problems, interventions, and evaluations. Here's a breakdown of its key features and considerations:
1.5K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

806
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
806
Nursing Clinical Information System01:27

Nursing Clinical Information System

881
Nursing Clinical Information System (NCIS)
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
881

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Privacy-preserving verification of preprocessing in federated learning for genomic data.

JAMIA open·2026
Same author

Explainable drug side effect prediction in central neural system via biologically informed graph neural network.

Translational psychiatry·2026
Same author

Privacy-Preserving Verification of ML Preprocessing via Model Behavior Indicators.

IEEE transactions on privacy·2026
Same author

Information extraction from clinical notes: are we ready to switch to large language models?

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

A medically grounded LLM agent-based tool to detect patient safety events in medical records: Identifying patient safety events with large language models.

medRxiv : the preprint server for health sciences·2025
Same author

Biomedical data repositories require governance for artificial intelligence/machine learning applications at every step.

JAMIA open·2025

Related Experiment Video

Updated: Sep 19, 2025

A Computer-Based Platform for Aiding Clinicians in Eating Disorder Analysis and Diagnosis
04:19

A Computer-Based Platform for Aiding Clinicians in Eating Disorder Analysis and Diagnosis

Published on: May 10, 2022

4.0K

Computerized diagnostic decision support systems-Isabel Pro versus ChatGPT-4 part II.

Joe M Bridges1, Xiaoqian Jiang1, Michael Ige1

  • 1D. Bradley McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, Houston, TX 77030, United States.

JAMIA Open
|June 17, 2025
PubMed
Summary

This study evaluated ChatGPT-4 for medical diagnosis, finding it improved some metrics but showed poor reproducibility and citation accuracy. Clinical use for diagnosis is currently limited due to these concerns.

Keywords:
ChatGPT-4Isabel Proartificial intelligencecomputer assisteddiagnosis

More Related Videos

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
05:33

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System

Published on: July 11, 2025

281
Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
07:51

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis

Published on: September 26, 2018

7.7K

Related Experiment Videos

Last Updated: Sep 19, 2025

A Computer-Based Platform for Aiding Clinicians in Eating Disorder Analysis and Diagnosis
04:19

A Computer-Based Platform for Aiding Clinicians in Eating Disorder Analysis and Diagnosis

Published on: May 10, 2022

4.0K
Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
05:33

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System

Published on: July 11, 2025

281
Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
07:51

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis

Published on: September 26, 2018

7.7K

Area of Science:

  • Medical Informatics
  • Artificial Intelligence in Healthcare
  • Clinical Decision Support Systems

Background:

  • Large language models (LLMs) like ChatGPT-4 are increasingly explored for clinical applications.
  • Evaluating the diagnostic accuracy and reliability of LLMs is crucial before widespread adoption.
  • Existing diagnostic decision support systems (DDSS) also have limitations that warrant comparison.

Purpose of the Study:

  • To assess the diagnostic accuracy of ChatGPT-4 compared to the Isabel Pro DDSS.
  • To investigate the impact of prompt engineering (Tree-of-Thought) and expert panel size on ChatGPT-4's performance.
  • To evaluate the reproducibility and reference citation accuracy of ChatGPT-4 in a diagnostic context.

Main Methods:

  • Utilized 201 cases from the New England Journal of Medicine for differential diagnosis generation.
  • Employed statistical measures including Mean Reciprocal Rank (MRR), Recall at Rank, and Average Rank.
  • Assessed reproducibility using r-squared calculations and evaluated citation accuracy for references and DOIs.

Main Results:

  • ChatGPT-4 improved MRR and Recall at 10 but had fewer correct diagnoses and lower average rank.
  • Revising the Isabel Pro differential improved Recall at 10 by 11%; a two-person expert panel yielded optimal results.
  • Reproducibility runs showed poor consistency (r-squared 0.34-0.44), and reference accuracy was low (34.8% for citations, 37.8% for DOIs).

Conclusions:

  • ChatGPT-4 demonstrates potential but faces significant challenges in diagnostic accuracy and reliability.
  • Poor reproducibility and a high frequency of fabricated references limit its current clinical utility for diagnosis.
  • Further development is required to address these critical issues before ChatGPT-4 can be safely integrated into clinical diagnostic workflows.