Related Experiment Video
Updated: Feb 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Size doesn't matter: Assessing the trustworthiness of large language models in medical contexts: A focus on epidural
Marina Del Barrio1, Kazim Laos2, María José Vilchez Lara3
1Hospital de Henares, Av. de Marie Curie, 0, Coslada, Madrid, 28822, Spain; Universidad Rey Juan Carlos, Av. de Atenas s/n, Alcorcon, Madrid, 28922, Spain.
Background:
Since the release of ChatGPT, numerous LLMs have emerged, providing easy access to information without the need for technical expertise. However, relying on these systems can influence important life decisions, such as the choice to use epidural analgesia during childbirth. Epidural analgesia is widely regarded as the "gold standard" for pain relief during childbirth. However, limited access to anaesthesiologists and gaps in knowledge may prompt individuals to seek information from unverified sources, including AI systems. Misinformation in this area can discourage the use of effective analgesia, highlighting the need to assess the accuracy of LLM-generated content.
Objective:
To evaluate the reliability of LLM-generated information regarding epidural analgesia.
Methods:
We posed 10 standardized questions about epidural analgesia to 12 LLMs, each question reformulated 10 times in both Spanish and English, resulting in 2400 responses. Two anaesthesiologists were involved in the assessment process. One expert performed the initial ratings, while the second independently verified the evaluations assessing the outputs using an extended SERVQUAL framework.
Results:
ChatGPT performed best, followed by Gemini 2. Medium-sized models, such as Phi-3 and OpenChat, outperformed several larger models like Llama-2 or Llama-3, challenging the notion that "bigger is better" and offering potential advantages in low-resource settings (e.g., Phi-3 outperformed Llama-2 with an average increase of 81% across all metrics). Specialized models did not show superior performance. Except for ChatGPT, English responses were generally more reliable than Spanish, with some Spanish outputs incoherent. ChatGPT also exhibited the least variability between responses.
Conclusions:
Despite promising performance, LLMs display limitations in medical contexts. Collaboration between national and international medical societies is crucial to develop evidence-based resources to guide LLM training and improve information trustworthiness.
Related Concept Videos
Local Anesthetics: Clinical Application as Epidural Anesthesia
Since epidural anesthetics can be infused through an epidural catheter, all types of drugs, including short-acting ones, can be administered. Chloroprocaine and lidocaine are examples of short and long-duration anesthetics, respectively. Bupivacaine...
Local Anesthetics: Clinical Application as Spinal Anesthesia
