Related Experiment Video
Updated: Sep 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Frontline Health Workers' Responses to Patient Inquiries With and Without Large Language Model Support in
Maria Moosa1, Oluwaseyi Malumi2, Solomon Chinedu1
1mDoc Healthcare, 1A Hakeem Dickson Drive Off T.F. Kuboye Street, Lekki Phase 1, Lagos, Lagos, NG.
Background:
Frontline health workers (FLWs) in low- and middle-income countries (LMICs) often face barriers that compromise care quality, including limited training, high patient loads, inadequate access to updated clinical guidance, and resource constraints. Working in underserved settings further exacerbates challenges to providing accurate, complete, and empathetic responses. Addressing these gaps requires innovative tools to support FLWs in real time.
Objective:
This study aims to evaluate the effectiveness of ChatGPT-3.5 in enhancing the quality of FLWs' responses to health inquiries in Lagos, Nigeria.
Methods:
In this observational study, 36 licensed FLWs (doctors, nurses, community health workers) practicing general or primary care in Lagos, Nigeria, generated responses to 15 patient-generated health questions covering maternal and neonatal health, family planning, and sexually transmitted infections. Each FLW produced responses using their own knowledge and available resources ("human-only") and with ChatGPT-3.5 assistance ("GPT-aided"). A panel of 4 clinicians evaluated responses using a 5-point Likert scale for accuracy, empathy, contextualization, completeness, safety, and overall preference. In addition, 8 non-clinician health seekers assessed a subset of responses across 3 themes: Caring, Trust, and Ease of Understanding.
Results:
Across 1080 responses from 36 FLWs, GPT-aided responses significantly outperformed human-only responses on all clinician-rated metrics (P<.001), with the largest gains observed in completeness, empathy, and overall preference, particularly among Community Health Extension Workers (CHEWs). Non-clinician evaluators also favored GPT-aided responses across all themes, with a strong preference for trust.
Conclusions:
This study provides empirical evidence that a large language model (ChatGPT-3.5) enhances the completeness, empathy, and perceived trust of responses by frontline health workers in Lagos, Nigeria. These findings underscore the potential of integrating ChatGPT-3.5 to strengthen healthcare delivery in resource-limited settings, highlighting the necessity for capacity building and further investigation into its safe, equitable, and effective implementation.