Related Experiment Video
Updated: Jun 17, 2025

A Competent Hepatocyte Model Examining Hepatitis B Virus Entry through Sodium Taurocholate Cotransporting Polypeptide as a Therapeutic Target
Published on: May 10, 2022
Assessing ChatGPT as a Medical Consultation Assistant for Chronic Hepatitis B: Cross-Language Study of English and
Yijie Wang1, Yining Chen2, Jifang Sheng1
1State Key Laboratory for Diagnosis and Treatment of Infectious Diseases, Collaborative Innovation Center for Diagnosis and Treatment of Infectious Disease, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China.
Insights
This study found that ChatGPT-4.0 offers improved medical consultation for chronic hepatitis B (CHB) patients compared to ChatGPT-3.5, especially in accuracy and information processing. However, both AI models require further training in emotional support for effective patient care.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing for Healthcare
- Digital Health and Patient Support
Background:
- Chronic hepatitis B (CHB) presents significant global health and economic challenges, particularly in resource-limited settings.
- Effective CHB management requires complex monitoring and patient adherence, often complicated by healthcare system constraints.
- Emerging AI assistants like ChatGPT offer potential solutions for improving CHB patient support and education.
Purpose of the Study:
- To evaluate the capabilities and limitations of ChatGPT-3.5 and ChatGPT-4.0 in providing personalized medical consultation for CHB patients.
- To assess AI performance across different linguistic contexts, specifically English and Chinese.
- To identify areas for improvement in AI-driven medical assistance for chronic disease management.
Main Methods:
- Compiled 96 questions from CHB guidelines, online communities, and search engines in English and Chinese.
- Presented questions to ChatGPT-3.5 and ChatGPT-4.0 in separate dialogues for response generation.
- Evaluated responses by senior physicians for informativeness, emotional management, consistency, and cautionary advice; used a true-or-false questionnaire to assess accuracy.
Main Results:
- ChatGPT-4.0 provided more comprehensive responses (74.5%) than ChatGPT-3.5 (61.6%).
- ChatGPT-4.0 demonstrated significantly higher accuracy (93.3%) on true-or-false questions compared to ChatGPT-3.5 (65.0%).
- Both models showed deficiencies in emotional management guidance, with ChatGPT-4.0 performing slightly better (8.1% vs. 3.2%).
Conclusions:
- ChatGPT models show basic utility as medical consultation assistants for CHB, with ChatGPT-4.0 offering enhanced capabilities.
- Language impacts ChatGPT-3.5 performance, while ChatGPT-4.0 mitigates this effect, highlighting the importance of model advancement.
- Both AI models require specific training in emotional management and disclaimer usage for effective clinical deployment in CHB care.
Background:
Chronic hepatitis B (CHB) imposes substantial economic and social burdens globally. The management of CHB involves intricate monitoring and adherence challenges, particularly in regions like China, where a high prevalence of CHB intersects with health care resource limitations. This study explores the potential of ChatGPT-3.5, an emerging artificial intelligence (AI) assistant, to address these complexities. With notable capabilities in medical education and practice, ChatGPT-3.5's role is examined in managing CHB, particularly in regions with distinct health care landscapes.
Objective:
This study aimed to uncover insights into ChatGPT-3.5's potential and limitations in delivering personalized medical consultation assistance for CHB patients across diverse linguistic contexts.
Methods:
Questions sourced from published guidelines, online CHB communities, and search engines in English and Chinese were refined, translated, and compiled into 96 inquiries. Subsequently, these questions were presented to both ChatGPT-3.5 and ChatGPT-4.0 in independent dialogues. The responses were then evaluated by senior physicians, focusing on informativeness, emotional management, consistency across repeated inquiries, and cautionary statements regarding medical advice. Additionally, a true-or-false questionnaire was employed to further discern the variance in information accuracy for closed questions between ChatGPT-3.5 and ChatGPT-4.0.
Results:
Over half of the responses (228/370, 61.6%) from ChatGPT-3.5 were considered comprehensive. In contrast, ChatGPT-4.0 exhibited a higher percentage at 74.5% (172/222; P<.001). Notably, superior performance was evident in English, particularly in terms of informativeness and consistency across repeated queries. However, deficiencies were identified in emotional management guidance, with only 3.2% (6/186) in ChatGPT-3.5 and 8.1% (15/154) in ChatGPT-4.0 (P=.04). ChatGPT-3.5 included a disclaimer in 10.8% (24/222) of responses, while ChatGPT-4.0 included a disclaimer in 13.1% (29/222) of responses (P=.46). When responding to true-or-false questions, ChatGPT-4.0 achieved an accuracy rate of 93.3% (168/180), significantly surpassing ChatGPT-3.5's accuracy rate of 65.0% (117/180) (P<.001).
Conclusions:
In this study, ChatGPT demonstrated basic capabilities as a medical consultation assistant for CHB management. The choice of working language for ChatGPT-3.5 was considered a potential factor influencing its performance, particularly in the use of terminology and colloquial language, and this potentially affects its applicability within specific target populations. However, as an updated model, ChatGPT-4.0 exhibits improved information processing capabilities, overcoming the language impact on information accuracy. This suggests that the implications of model advancement on applications need to be considered when selecting large language models as medical consultation assistants. Given that both models performed inadequately in emotional guidance management, this study highlights the importance of providing specific language training and emotional management strategies when deploying ChatGPT for medical purposes. Furthermore, the tendency of these models to use disclaimers in conversations should be further investigated to understand the impact on patients' experiences in practical applications.

