Related Experiment Video
Updated: Jun 13, 2025

Microscopic Electric Rotary Grinding of Plaques Combined with Graft Repair in the Management of Peyronie's Disease
Published on: March 15, 2024
Prompt matters: evaluation of large language model chatbot responses related to Peyronie's disease
Christopher J Warren1, Victoria S Edmonds1, Nicolette G Payne1
1Department of Urology, Mayo Clinic Arizona, Phoenix, AZ 85054, United States.
Introduction:
Despite direct access to clinicians through the electronic health record, patients are increasingly turning to the internet for information related to their health, especially with sensitive urologic conditions such as Peyronie's disease (PD). Large language model (LLM) chatbots are a form of artificial intelligence that rely on user prompts to mimic conversation, and they have shown remarkable capabilities. The conversational nature of these chatbots has the potential to answer patient questions related to PD; however, the accuracy, comprehensiveness, and readability of these LLMs related to PD remain unknown.
Aims:
To assess the quality and readability of information generated from 4 LLMs with searches related to PD; to see if users could improve responses; and to assess the accuracy, completeness, and readability of responses to artificial preoperative patient questions sent through the electronic health record prior to undergoing PD surgery.
Methods:
The National Institutes of Health's frequently asked questions related to PD were entered into 4 LLMs, unprompted and prompted. The responses were evaluated for overall quality by the previously validated DISCERN questionnaire. Accuracy and completeness of LLM responses to 11 presurgical patient messages were evaluated with previously accepted Likert scales. All evaluations were performed by 3 independent reviewers in October 2023, and all reviews were repeated in April 2024. Descriptive statistics and analysis were performed.
Results:
Without prompting, the quality of information was moderate across all LLMs but improved to high quality with prompting. LLMs were accurate and complete, with an average score of 5.5 of 6.0 (SD, 0.8) and 2.8 of 3.0 (SD, 0.4), respectively. The average Flesch-Kincaid reading level was grade 12.9 (SD, 2.1). Chatbots were unable to communicate at a grade 8 reading level when prompted, and their citations were appropriate only 42.5% of the time.
Conclusion:
LLMs may become a valuable tool for patient education for PD, but they currently rely on clinical context and appropriate prompting by humans to be useful. Unfortunately, their prerequisite reading level remains higher than that of the average patient, and their citations cannot be trusted. However, given their increasing uptake and accessibility, patients and physicians should be educated on how to interact with these LLMs to elicit the most appropriate responses. In the future, LLMs may reduce burnout by helping physicians respond to patient messages.
Related Concept Videos
Disorders of the Male Reproductive System
Prostate disorders are another major concern. These conditions can impair urinary flow due to the prostate's location around the urethra....
Treatment for Pulmonary Arterial Hypertension: Phosphodiesterase Inhibitors
Among the PDE5 inhibitors, sildenafil (Revatio) stands out as a competitive and selective inhibitor. It operates by elevating cellular levels of cGMP and augmenting signaling through the cGMP-PKG pathway, promoting vasodilation. Upon oral...
Male Sexual Response: Erection & Ejaculation
The blood filling the erectile tissues compresses the veins, which helps to prevent blood from leaving...

