Related Experiment Video
Updated: Jul 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Controlled Benchmark Evaluation of a Geographically Grounded Digital Health Chatbot Against a General-Purpose Model
Zhaoqiang Zhou1, John Geracitano1, Saif Khairat1
1University of North Carolina at Chapel Hill, NC, USA.
Abstract:
Large language models are increasingly used for health information delivery, but their value for geographically grounded digital health information services remains unclear. We developed a domain-specific, LLaMA-based Digital Health Index (DHI) chatbot grounded in digital health datasets and compared it with GPT-5.3 in a controlled benchmark evaluation. The benchmark included 15 items spanning five scenario types and three difficulty levels. Responses were independently rated by five health informatics experts using a multidimensional 5-point Likert rubric. Overall, the DHI chatbot outperformed GPT-5.3 across all six core domains, with the largest gains in geographic handling, evidence transparency, and accuracy. These advantages remained statistically significant after false discovery rate correction and were broadly consistent across difficulty levels. The findings suggest that geographically grounded, domain-specific conversational AI may better support accurate, transparent, and locally relevant digital health information services than general-purpose models.