Related Experiment Video
Updated: Jan 7, 2026

Rapid Detection of Fecal Antigen of Helicobacter pylori Infection Based on Double Antibody Sandwich Detection Technology
Published on: May 23, 2025
Performance of Large Language Models in Chinese Language Medical Counseling on Helicobacter pylori
Mingjun Zhang1, Shiming Zhou1, Shulin Zhang1
1Department of Gastroenterology, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University, Beijing, People's Republic of China.
Background:
H. pylori infection is a worldwide health issue, fueling rising demand for medical counseling. LLMs have the potential to serve in medical counseling. However, their performance remains unclear.
Objective:
This study aimed to evaluate the effectiveness of LLMs in providing H. pylori related medical counseling in a Chinese context.
Methods:
20 H. pylori-related questions were collected, covering four domains: definition and symptoms, diagnosis, treatment, and prevention. Each question was asked thrice in Chinese to each LLM. We assessed the responses across five dimensions (accuracy, relevance, completeness, clarity, and reliability).
Results:
1. In the first batch of tests, the overall performance distribution was 33.3% good, 66.1% medium, and 0.6% poor, respectively. No significant differences were observed among the three LLMs (p=0.158). Good performance was observed with 47.8% in accuracy, 53.9% in relevance, 68.3% in completeness, 36.7% in clarity, and 36.1% in reliability. No significant differences were observed in accuracy, relevance, completeness, or clarity. Reliability differed significantly (p<0.001), with Ernie Bot achieving the best performance. 2. The second test batch yielded performance rates of 70.6% good, 29.4% medium, and 0% poor, with a significant difference among the three LLMs (p=0.018). Doubao attained the best performance, surpassing other models in relevance and clarity. 3. The newly assessed AI batch showed markedly superior overall performance to the counterpart evaluated more than a year prior.
Conclusion:
This study is the first to evaluate the effectiveness of various LLMs in H. pylori-related medical counseling in a real-world setting. The study showed that while LLMs generally performed acceptably in terms of accuracy, relevance, and completeness, their clarity and reliability were less satisfactory. Ernie Bot, developed by Chinese company, outperformed ChatGPT in certain aspects of medical counseling in Chinese. With the guidance of professionals, LLMs can serve as potential aids for medical counseling.
Related Concept Videos
Treating Helicobacter pylori in Peptic Ulcers: Antimicrobial Therapy
Peptic Ulcer Disease III: Clinical Manifestations and Diagnostic Studies
Few clinical manifestations differentiate gastric ulcers from duodenal ulcers. Distinctions in the location, timing, and pain relief are crucial for healthcare providers in differentiating between gastric and duodenal ulcers during clinical assessments.

