Related Experiment Video
Updated: May 23, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Medical Misinformation in AI-Assisted Self-Diagnosis: Development of a Method (EvalPrompt) for Analyzing Large
Troy Zada1, Natalie Tam1, Francois Barnard1
1Department of Management Sciences and Engineering, University of Waterloo, 200 University Avenue West, Waterloo, ON, N2L 3G1, Canada, 1 5198884567 ext 33279.
Large language models (LLMs) like ChatGPT show modest capabilities for self-diagnosis, with accuracy below passing thresholds. Relying solely on LLMs for medical information risks misinformation due to unclear and inaccurate responses.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) are rapidly integrating into healthcare, prompting discussions on their potential to improve quality and accessibility.
- While LLMs can pass medical exams, their use for self-diagnosis and potential to spread misinformation requires evaluation.
Purpose of the Study:
- To assess the effectiveness of LLMs, specifically ChatGPT, for individual self-diagnosis.
- To evaluate the clarity, correctness, and robustness of LLM responses in simulated self-diagnosis scenarios.
Main Methods:
- Developed the EvalPrompt methodology using medical licensing examination questions.
- Experiment 1: Prompted ChatGPT with open-ended questions simulating self-diagnosis.
- Experiment 2: Assessed robustness by using sentence dropout in responses to mimic missing information.
Main Results:
- ChatGPT-4.0 responses were deemed correct for only 31% of questions in Experiment 1.
- In Experiment 2, 61% of responses remained correct, indicating robustness but also potential for misinformation.
- ChatGPT-4.0 fell below the 60% passing threshold for clarity and correctness.
Conclusions:
- LLM capabilities for self-diagnosis are currently modest, with responses often unclear and inaccurate.
- Caution is advised when using LLM-generated medical advice due to significant misinformation risks.
- Development of a comprehensive self-diagnosis dataset is needed to improve LLM reliability in healthcare.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
MALDI-TOF Mass Spectrometry
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Enzyme-Linked Immunosorbent Assay
There are many different types of ELISAs, but they all involve an antibody molecule whose constant region binds an enzyme, leaving the variable region free to bind its specific antigen. Enzyme-substrate reaction allows the antigen to be visualized or...

