Related Experiment Video
Updated: May 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
An automated framework for assessing how well LLMs cite relevant medical references
Kevin Wu1, Eric Wu2, Kevin Wei3
1Department of Biomedical Data Science, Stanford University, Stanford, CA, USA.
Large language models (LLMs) often fail to support health claims with cited sources. A new study reveals 50-90% of LLM responses lack adequate source support, impacting medical query trustworthiness.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly utilized for health-related inquiries.
- Ensuring LLM-generated health information is supported by credible references is critical.
- The degree to which LLM-cited sources actually support their claims is not well understood.
Purpose of the Study:
- To introduce SourceCheckup, an automated pipeline for evaluating source relevance and support in LLM responses.
- To assess the reliability of sources cited by popular LLMs for medical queries.
- To quantify the extent of source support for LLM-generated medical statements.
Main Methods:
- Developed an agent-based pipeline named SourceCheckup.
- Evaluated seven popular LLMs using a dataset of 800 medical questions.
- Analyzed 58,000 statement-source pairs from LLM responses to common medical queries.
Main Results:
- Between 50% and 90% of LLM responses were not fully supported or were contradicted by their cited sources.
- Even advanced models like GPT-4o with web search showed approximately 30% of statements unsupported.
- Independent medical professional assessments corroborated the findings of inadequate source support.
Conclusions:
- Current LLMs exhibit significant limitations in providing trustworthy medical references.
- A substantial portion of LLM-generated health information lacks adequate evidential support.
- Further development is needed to enhance the reliability of LLMs for health-related applications.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Health Literacy
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Assessing Blood pressure in the Leg
Preparation:
Improving Translational Accuracy
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...