Related Experiment Video
Updated: Jul 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reliability and readability of adenoid hypertrophy information generated by five publicly accessible LLM chatbots: A
Xiaoming Qian1, Zhishui Wu2, Jing Li3
1Department of Otolaryngology, The Third Affiliated Hospital of Zhengzhou University, Zhengzhou, Henan, China.
Large language models (LLMs) show an imbalance in generating reliable yet easy-to-read patient information on adenoid hypertrophy. While Perplexity offered higher quality, Gemini provided more readable responses, but neither met ideal readability standards.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Patient Education
Background:
- Adenoid hypertrophy information is crucial for patient understanding and care.
- Large Language Models (LLMs) are increasingly used for generating health information.
- Evaluating the quality of LLM-generated health content is essential.
Purpose of the Study:
- To assess the performance of five major LLM chatbots in generating patient-oriented information on adenoid hypertrophy.
- To evaluate the reliability and readability of LLM-generated content for a common pediatric condition.
Main Methods:
- Collected 63 FAQs on adenoid hypertrophy covering etiology, symptoms, and treatment.
- Submitted questions to five LLMs (Perplexity, Copilot, ChatGPT, DeepSeek, Gemini) via official interfaces.
- Assessed reliability using DISCERN, EQIP, JAMA benchmarks, and GQS; measured readability with six indices; clinicians scored responses.
Main Results:
- Significant differences in reliability were observed among LLMs (P <0.001).
- Perplexity and Copilot performed best on reliability benchmarks (DISCERN, EQIP, JAMA).
- No LLM met the recommended sixth-grade reading level; Gemini offered the best readability (FRES: 61.95±9.64).
Conclusions:
- LLMs exhibit a trade-off between reliability and readability for adenoid hypertrophy information.
- Perplexity excelled in information quality, while Gemini provided more accessible text.
- Future LLM development should prioritize source transparency and text simplification for improved AI-assisted health communication.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:14Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025