Related Experiment Video
Updated: Jun 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models in Answering Healthcare Delivery Questions: A Quantitative Cross-Sectional Study
Mohsen Khosravi1, Zahra Zamaninasab2, Fatemeh Khosravi3
1Social Determinants of Health Research Center Birjand University of Medical Sciences Birjand Iran.
Background And Aims:
The use of Large Language Model (LLM)-based chatbots across various fields has yielded positive outcomes. Understanding the health service delivery system offers numerous benefits. This study aimed to analyze the performance of LLMs in answering healthcare delivery questions.
Methods:
A validated questionnaire relevant to the research context was administered to a sample of LLM-based chatbots. The chatbots evaluated in this study included GPT-4.1-mini, Gemini 2.5, Copilot 2025, and Perplexity. A written prompt was provided to facilitate response generation by the chatbots. To analyze and compare the performance of the AI models in addressing the research questions, confusion matrices were constructed, and key metrics-sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy-were calculated.
Results:
The initial assessment of the chatbots showed perfect sensitivity (1.00), accurately identifying all true positives without false negatives. Specificity varied, with ChatGPT and Perplexity at 0.50, Gemini at 0.43, and Copilot at 0.33. Positive predictive values (PPV) ranged from 0.67 (Gemini) to 0.75 (ChatGPT and Perplexity), while negative predictive values (NPV) were uniformly perfect (1.00). Overall accuracy was highest for ChatGPT and Perplexity (0.80), with Gemini and Copilot at 0.73. In the second round, sensitivity remained perfect for all chatbots. Gemini achieved the highest specificity (0.80), followed by ChatGPT (0.67), Perplexity (0.60), and Copilot (0.50). PPVs improved, ranging from 0.75 (Copilot) to 0.91 (Gemini). NPVs remained perfect (1.00) across all models. Overall accuracy led by Gemini (0.93), with ChatGPT and Perplexity both at 0.87, and Copilot at 0.80.
Conclusion:
ChatGPT and Perplexity showed the highest initial performance, while the second round revealed improvements in most chatbots, especially in specificity and accuracy, with Gemini performing best. Further research is needed for deeper insights.
