Related Experiment Video
Updated: Aug 6, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
Evaluating the reliability and information quality of ChatGPT responses on vaccine hesitancy: an expert panel study
Hande Can Güvercin1, Cemal Koçak2, Mehmet Furkan Aytekin1
1Department of Public Health, Ankara University Faculty of Medicine, Ankara, Türkiye.
Background:
Vaccine hesitancy is a major public health problem that threatens global immunisation programmes. The use of artificial intelligence-based chatbots to access health information is increasing. This study aimed to evaluate the responses of ChatGPT-4o to frequently asked questions about vaccine hesitancy in terms of health information quality, based on expert opinion.
Methods:
This cross-sectional expert evaluation study included 20 questions on vaccine hesitancy. These questions were submitted to ChatGPT-4o, and the responses were evaluated by nine experts from the fields of public health, paediatrics, and infectious diseases. The evaluation was based on six criteria synthesised from HONcode, DISCERN, JAMA Benchmarks, CRAAP, GQS, and QUEST: scientific accuracy, comprehensiveness, understandability, correction of misinformation, source attribution, and actionability. Each criterion was scored on a scale from 1 to 10. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC) and, as a prevalence-robust sensitivity analysis, Gwet's AC2 with quadratic weights.
Results:
The highest mean scores were found for understandability (9.4 ± 1.1) and scientific accuracy (8.9 ± 1.3), while the lowest mean score was found for source attribution (4.5 ± 3.4). At the question level, the highest scores were obtained for Question 4, on aluminium in vaccines, and Question 14, on claims that pharmaceutical companies endanger children's health by producing and promoting vaccines. The lowest scores were obtained for Question 5, on live vaccines during pregnancy, and Question 6, on natural immunity. ICC analysis showed moderate agreement only for source attribution (ICC = 0.582; p < 0.001); the low ICC values for the other criteria were attributable to a ceiling effect, as confirmed by Gwet's AC2, which indicated substantially higher agreement for the high-scoring criteria.
Conclusions:
ChatGPT-4o can generate scientifically accurate and understandable responses to questions about vaccine hesitancy. However, it shows a systematic shortcoming in source attribution. Although the model has potential as a supportive tool in public health communication, its outputs should be reviewed by experts and supported with verifiable sources.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Reliability and Validity
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Surveys
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
The Availability Heuristic