Related Experiment Video
Updated: Jul 1, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of large language models as an information resource on functional hypothalamic amenorrhea for patients
Nancy Safwan1,2, Jana Karam1,2, Sarah L Berga1,2,3,4
1Division of General Internal Medicine, Mayo Clinic, Jacksonville, FL, United States.
Introduction:
To assess and compare the accuracy, readability, and overall performance of large language models (LLMs) in answering questions about functional hypothalamic amenorrhea (FHA) for patients and healthcare professionals.
Methods:
A total of 11 patient-level and 15 clinician-level FHA-related questions were entered separately into four LLMs: ChatGPT 3.5 (free version), ChatGPT 4.0 (updated, paid subscription), Gemini, and OpenEvidence. OpenEvidence was used only for clinician-based questions. Responses were evaluated by three expert reviewers blinded to the LLM used who rated them as accurate and complete, accurate but incomplete, or inaccurate. A fourth reviewer resolved discordant scores. Readability for patient-level questions was assessed using the Flesch Reading Ease Score (FRES) and word count. Lower FRES scores indicate more difficult reading. Accuracy and completeness were compared using odds ratios (95% CI) with ChatGPT 3.5 as the reference model, and differences in readability were analyzed using Friedman's test.
Results:
LLM performance varied across question types. For patient-level questions, ChatGPT 4.0 achieved the highest accuracy (9 of 11; 82%), followed by ChatGPT 3.5 and Gemini (each 8 of 11; 73%), with no statistically significant differences. Among clinician-level questions, OpenEvidence demonstrated perfect accuracy (15 of 15; 100%), compared with 93% for and 80% for ChatGPT 4.0 and Gemini. Completeness followed similar patterns, with OpenEvidence providing the most complete clinician responses (93%) and ChatGPT 4.0 the most complete patient-level responses (89%). Readability differed significantly among models (p = 0.012), with Gemini producing the most readable patient-level content (median FRES 43.5 [IQR 36.8-53.4]) compared with ChatGPT 3.5 (30.6 [16.8-48.4]) and ChatGPT 4.0 (28.8 [22.1-37.6]). Word counts did not differ significantly (p = 0.39).
Discussion:
LLMs demonstrated good overall performance in answering FHA-related questions but often provided incorrect or incomplete information. Fine tuning field-specific data, engineered prompts, and obtaining human-in-the-loop feedback may help improve the accuracy of these models.
Related Concept Videos
Hormonal Regulation
Hormonal Regulation
Hypothalamic-Pituitary Axis
Clearance Models: Physiological Models
The organ's clearance rate depends on the blood flow to the organ and the extraction ratio (E). The extraction ratio describes the organ's proficiency in drug...
Introduction to Language of Pathophysiology ll
Hormones Regulating Blood Glucose
In addition to accelerating glucose uptake and utilization, insulin has...
