Related Experiment Videos
Mapping applications and evaluations of LLM-enabled AI chatbots for health purposes: a scoping review
Yuan Wang1, Romy Rw2, Yanjiao Deng3
1Wee Kim Wee School of Communication and Information, Nanyang Technological University, Singapore, Singapore.
Background:
Large Language Model (LLM)-enabled artificial intelligence (AI) chatbots are increasingly shaping health communication by mediating how patients, health professionals, researchers, and health institutions seek, interpret, produce, and act on health information. Existing reviews have largely focused on a single stakeholder group or clinical domain, leaving unclear how these systems are applied and evaluated across stakeholder groups and health purposes.
Objective:
This scoping review mapped empirical studies of LLM-enabled AI chatbot applications and evaluations for health purposes across four stakeholder groups: the public or patients, health professionals, health researchers and students, and health institutions. We characterized the purposes and evidence patterns associated with these applications and evaluations.
Methods:
We searched nine databases for peer-reviewed, English-language empirical studies available through July 2025. After screening, 286 articles were coded for study characteristics, methodology, evidence type, chatbot modality, LLM type, health topic, stakeholder group, and health purpose.
Results:
The included articles were published between 2023 and 2025. Most examined general-purpose LLMs (n = 267, 93.4%). Studies were concentrated in the United States and China and primarily examined text-based chatbots. Noncommunicable or chronic diseases and general health information were the most common health topics. Ten health-related purposes were identified across four stakeholder groups and two umbrella categories. Public-facing applications involved the general public or patients and encompassed four purposes: (1) seeking health information, (2) symptom assessment and self-care management, (3) emotional support, and (4) preventive and transitional care support. Professional and organizational applications involved health professionals, health researchers and students, and health institutions, and encompassed six purposes: (5) clinical decision support, (6) improving administrative efficiency, (7) supporting patient interaction and doctor-patient communication, (8) enhancing continuing education, (9) facilitating health research, and (10) institutional workflow support and care navigation.
Conclusion:
This review provides a stakeholder-purpose mapping of AI chatbot applications and evaluations in health care. The evidence base is concentrated in text-based, lower-acuity, and information-oriented contexts and derives primarily from output evaluations and self-reported perceptions or use. Key gaps concern institutional integration and governance, higher-stakes and longitudinal contexts, and evidence across populations, languages, and interaction modalities.