Related Experiment Video
Updated: Jan 21, 2026

Comparing Objective Conjunctival Hyperemia Grading and the Ocular Surface Disease Index Score in Dry Eye Syndrome During COVID-19
Published on: May 25, 2022
Estimation of IPSS and OABSS scores using ChatGPT-4o: a comparative validation study in Korea
Hoyoung Bae1,2, Gyu Min Lee1,2, Jiehyeon Lee1,2
1Department of Urology, Seoul National University Boramae Medical Center, Sindaebang 2(i)-dong, Dongjak-gu, Seoul, 07061, Korea.
ChatGPT-4o demonstrated moderate accuracy in estimating International Prostate Symptom Score (IPSS) and Overactive Bladder Symptom Score (OABSS), showing potential to aid in clinical data collection.
Area of Science:
- Artificial Intelligence in Healthcare
- Urology
- Clinical Informatics
Background:
- Patient-reported outcome measures (PROMs) like the International Prostate Symptom Score (IPSS) and Overactive Bladder Symptom Score (OABSS) are crucial for diagnosing and managing lower urinary tract symptoms (LUTS) and overactive bladder (OAB).
- Accurate symptom scoring relies on patient recall and questionnaire completion, which can be subject to errors or omissions.
- Large language models (LLMs) like ChatGPT-4o offer a potential new avenue for data extraction and analysis from clinical notes.
Purpose of the Study:
- To assess the performance of ChatGPT-4o in estimating IPSS and OABSS using natural language descriptions and full outpatient records.
- To compare AI-generated scores with actual patient-reported questionnaire scores.
Main Methods:
- A study involving 91 patients, with 52 completing IPSS and 77 completing OABSS.
- ChatGPT-4o processed verbatim patient symptom statements and urologist-written medical records.
- Statistical analyses included paired t-tests, Cohen's kappa, Spearman's correlation, Bland-Altman plots, McNemar's test, and ROC curve analysis for score and diagnostic classification comparison.
Main Results:
- ChatGPT-4o underestimated mean IPSS (11.2 vs. 13.6, p=0.006) but showed no significant difference for OABSS (6.99 vs. 6.86, p=0.686).
- High diagnostic agreement was observed for LUTS (AUC 0.81) and OAB (AUC 0.91), with good concordance in quality of life and urgency incontinence.
- Spearman's correlation coefficients were 0.60 for IPSS and 0.70 for OABSS, indicating moderate to strong correlations.
Conclusions:
- ChatGPT-4o achieved moderate, clinically acceptable accuracy in estimating IPSS and OABSS.
- The AI's diagnostic classification performance was comparable to actual scores, especially for OABSS and quality of life.
- ChatGPT-4o shows promise as a complementary tool for traditional questionnaires, particularly when patient data is incomplete.
More Related Videos
15:49Flexible Colonoscopy in Mice to Evaluate the Severity of Colitis and Colorectal Tumors Using a Validated Endoscopic Scoring System
Published on: October 16, 2013
06:14In Vivo Protocol of Controlled Subconcussive Head Impacts for the Validation of Field Study Data
Published on: April 18, 2019
Related Concept Videos
Reliability and Validity
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Introduction to z Scores
z scores...
Introduction to z Scores
z scores...
z Scores and Area Under the Curve
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...