Related Experiment Video
Updated: Jul 2, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy and Readability of Kidney Stone Patient Information Materials Generated by a Large Language Model Compared
Abdulghafour Halawani1, Alec Mitchell2, Mohammadali Saffarzadeh2
1Department of Urology, King Abdulaziz University, Jeddah, Saudi Arabia; Department of Urological Sciences, University of British Columbia, Stone Centre at Vancouver General Hospital, Vancouver, British Columbia, Canada.
Objective:
To compare the readability and accuracy of large language model generated patient information materials (PIMs) to those supplied by the American Urological Association (AUA), Canadian Urological Association (CUA), and European Association of Urology (EAU) for kidney stones.
Methods:
PIMs from AUA, CUA, and EAU related to nephrolithiasis were obtained and categorized. The most frequent patient questions related to kidney stones were identified from an internet query and input into GPT-3.5 and GPT-4. PIMs and ChatGPT outputs were assessed for accuracy and readability using previously published indexes. We also assessed changes in ChatGPT outputs when a reading level was specified (grade 6).
Results:
Readability scores were better for PIMs from the CUA (grade level 10-12), AUA (8-10), or EAU (9-11) compared to the chatbot. GPT-3.5 had the worst readability scores at grade 13-14 and GPT-4 was likewise less readable than urologic organization PIMs with scores of 11-13. While organizational PIMs were deemed to be accurate, the chatbot had high accuracy with minor details omitted. GPT-4 was more accurate in general stone information, dietary and medical management of kidney stones topics in comparison to GPT-3.5, while both models had the same accuracy in the surgical management of nephrolithiasis topics.
Conclusion:
Current PIMs from major urologic organizations for kidney stones remain more readable than publicly available GPT outputs, but they are still higher than the reading ability of the general population. Of the available PIMs for kidney stones, those from the AUA are the most readable. Although Chatbot outputs for common kidney stone patient queries have a high degree of accuracy with minor omitted details, it is important for clinicians to understand their strengths and limitations.
Related Concept Videos
Antihypertensive Drugs: Potassium-Sparing Diuretics
Antihypertensive Drugs: Direct Renin Inhibitors
Internal Anatomy of the Kidney
Anatomical Position and Dimensions
The kidneys are retroperitoneal organs positioned against the posterior abdominal wall on either side of the spine, roughly between the twelfth thoracic and third lumbar vertebrae. Each kidney is typically 10-12 cm long, 5-6 cm wide, and 3-4 cm thick, weighing about 150 grams.
Renal Cortex
The outermost region of the kidney is the...
Drug Elimination by Renal Route: Glomerular Filtration
Heart Failure Drugs: Inhibitors of Renin-Angiotensin System
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.

