Related Experiment Video
Updated: Jul 13, 2026

Application of the Intelligent High-Throughput Antimicrobial Sensitivity Testing/Phage Screening System and Lar Index of Antimicrobial Resistance
Published on: July 21, 2023
Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment
Rong Liu1, Ying Chen2, Wenzhuo Zhao3
1Department of Pharmacy, The Affiliated Changsha Central Hospital, Hengyang Medical School, University of South China, Changsha, People's Republic of China.
Background:
Pulmonary tuberculosis (TB) is a chronic infectious disease that burdens patients and public health systems. Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust. Large language models (LLMs) show promise for TB health education, but systematic evaluation is lacking.
Objective:
To evaluate five large language models in pulmonary tuberculosis Q&A (Question and Answer) scenarios and examine the effects of different large language models and TB health education themes on response quality, readability, and reliability, thereby supporting standardized artificial intelligence (AI)-assisted health education.
Methods:
This cross-sectional study was conducted from October 5 to October 11, 2025. Twenty pulmonary tuberculosis-related questions were developed with the assistance of a respiratory physician and categorized into five themes. The questions were entered into five large language models (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT) to generate 100 text responses. Quality was evaluated using patient-education suitability (C-PEMAT-P) and Global Quality Score (GQS), and readability using seven indices, including Automated Readability Index (ARI) and Flesch Reading Ease Score (FRES). Statistical analyses included One-way analysis of variance (One-way ANOVA), Kruskal-Wallis, and correlation analysis.
Results:
GPT-5 achieved the highest C-PEMAT-P scores, followed by Doubao, while GQS scores were similar across models. Models showed significant differences on several readability indices, whereas themes had limited effects. Quality indicators were modestly associated with readability, while readability indices were strongly intercorrelated.
Conclusions:
Model type is a key determinant of TB health education text quality. Quality and reading difficulty are related but relatively independent and should be jointly considered when selecting large language model (LLM)-generated materials. Further studies should include more models, diseases, and patient-reported outcomes to optimize AI-assisted health education.
Related Concept Videos
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Improving Translational Accuracy
