Related Experiment Video
Updated: Aug 5, 2026

E-Patient Counseling Trial (E-PACO): Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Measuring Consistency Between Large Language Models' Responses to Preventive Care Queries and Official US Preventive
Tim Johnson1, Wolfgang Gaissmaier2
1Atkinson Graduate School of Management, Willamette University, Salem, OR, United States.
Large language models (LLMs) show varying consistency with US Preventive Services Task Force (USPSTF) recommendations for preventive care. Newer models and specific prompting strategies significantly improve LLM accuracy in providing evidence-based guidance.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
- Public Health Informatics
Background:
- Large language models (LLMs) offer potential for scalable, individualized preventive care guidance.
- Previous research indicated mixed performance of limited LLMs on preventive care topics.
- A broader evaluation of diverse LLMs across a wider range of preventive care is necessary.
Purpose of the Study:
- To assess the consistency of popular large language models (LLMs) with US Preventive Services Task Force (USPSTF) recommendations.
- To evaluate LLM performance across a comprehensive set of USPSTF preventive care guidelines.
Main Methods:
- 35 LLMs were evaluated against 142 USPSTF recommendations (as of May 2025) in two study waves.
- Simulated user queries were used to assess LLM responses for concordance with USPSTF guidelines.
- Advanced prompting techniques (chain-of-thought, few-shot, role-based, iterative) were tested to optimize LLM performance.
Main Results:
- LLM-USPSTF concordance varied, with the best models achieving 66.92% (Wave 1) and 87.77% (Wave 2) baseline concordance.
- Chain-of-thought and role-based prompting significantly improved LLM accuracy, reaching up to 95.96% concordance.
- LLMs frequently avoided definitive recommendations, a primary source of discrepancy; however, grading tasks showed high accuracy (up to 94.37%).
Conclusions:
- LLM performance in aligning with USPSTF preventive care recommendations is inconsistent but improving.
- LLMs' tendency to avoid definitive statements impacts accuracy, necessitating careful prompt engineering.
- Advanced prompting strategies enhance LLM concordance, highlighting their potential for reliable clinical decision support.
Related Concept Videos
Preventive Healthcare Services
Healthcare Associated Infections II: Preventive Measures
The best practices for preventing healthcare-associated infections include hand hygiene, patient risk...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Levels of Health Promotion and Illness Prevention
In primary prevention, actions taken before disease onset prevent the disease from...