Related Experiment Video
Updated: Sep 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based
Qiqi Zheng1, Ru Chen1, Mingming Cai1
1Nursing Unit, Department of Infectious Diseases and Hepatology Center, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, China.
Background:
Search-enabled large language model interfaces are increasingly used by the public for health information, but their performance in mpox-related public health consultation remains unclear. This study evaluated their safety, accuracy, empathy, reliability/information quality, and readability.
Methods:
We conducted a single-query comparative cross-sectional evaluation using 52 predefined mpox-related public consultation questions. Each question was submitted once to each of six search-enabled LLM interfaces, yielding 312 first responses. Responses were assessed against a guideline-based reference framework. Safety was coded as a binary outcome, while accuracy and empathy were rated on 5-point scales. Reliability/information quality was evaluated using DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was assessed using six established readability indices. Five trained raters independently evaluated the human-scored outcomes.
Results:
Unsafe responses were relatively infrequent but occurred in all six interfaces, with safe-response rates ranging from 86.5 to 92.3%. No pairwise difference in Safety remained statistically significant after Benjamini-Hochberg correction. Overall differences across interfaces were statistically significant for Accuracy, Empathy, all four reliability/information quality measures, and all six readability indices. Benjamini-Hochberg-adjusted post hoc analyses identified outcome-specific pairwise differences, although the pairwise patterns varied across measures.
Conclusion:
The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability. Although unsafe responses were relatively uncommon, potentially harmful outputs occurred in every interface. These findings support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation. The results represent a time- and configuration-specific interface-level snapshot; they should not be attributed to the underlying base models in isolation or interpreted as establishing reproducible performance or a stable hierarchy across sessions, versions, or settings.