Assessing ChatGPT as a Medical Consultation Assistant for Chronic Hepatitis B: Cross-Language Study of English and

Yijie Wang1, Yining Chen2, Jifang Sheng1

  • 1State Key Laboratory for Diagnosis and Treatment of Infectious Diseases, Collaborative Innovation Center for Diagnosis and Treatment of Infectious Disease, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China.

PubMed

Insights

This study found that ChatGPT-4.0 offers improved medical consultation for chronic hepatitis B (CHB) patients compared to ChatGPT-3.5, especially in accuracy and information processing. However, both AI models require further training in emotional support for effective patient care.

Area of Science:

  • Artificial Intelligence in Medicine
  • Natural Language Processing for Healthcare
  • Digital Health and Patient Support

Background:

  • Chronic hepatitis B (CHB) presents significant global health and economic challenges, particularly in resource-limited settings.
  • Effective CHB management requires complex monitoring and patient adherence, often complicated by healthcare system constraints.
  • Emerging AI assistants like ChatGPT offer potential solutions for improving CHB patient support and education.

Purpose of the Study:

  • To evaluate the capabilities and limitations of ChatGPT-3.5 and ChatGPT-4.0 in providing personalized medical consultation for CHB patients.
  • To assess AI performance across different linguistic contexts, specifically English and Chinese.
  • To identify areas for improvement in AI-driven medical assistance for chronic disease management.

Main Methods:

  • Compiled 96 questions from CHB guidelines, online communities, and search engines in English and Chinese.
  • Presented questions to ChatGPT-3.5 and ChatGPT-4.0 in separate dialogues for response generation.
  • Evaluated responses by senior physicians for informativeness, emotional management, consistency, and cautionary advice; used a true-or-false questionnaire to assess accuracy.

Main Results:

  • ChatGPT-4.0 provided more comprehensive responses (74.5%) than ChatGPT-3.5 (61.6%).
  • ChatGPT-4.0 demonstrated significantly higher accuracy (93.3%) on true-or-false questions compared to ChatGPT-3.5 (65.0%).
  • Both models showed deficiencies in emotional management guidance, with ChatGPT-4.0 performing slightly better (8.1% vs. 3.2%).

Conclusions:

  • ChatGPT models show basic utility as medical consultation assistants for CHB, with ChatGPT-4.0 offering enhanced capabilities.
  • Language impacts ChatGPT-3.5 performance, while ChatGPT-4.0 mitigates this effect, highlighting the importance of model advancement.
  • Both AI models require specific training in emotional management and disclaimer usage for effective clinical deployment in CHB care.
Abstract