Related Experiment Video
Updated: Jun 11, 2025

06:16
Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
3.7K
Optimizing ChatGPT's Interpretation and Reporting of Delirium Assessment Outcomes: Exploratory Study
Yong K Choi1, Shih-Yin Lin2, Donna Marie Fick3
1Department of Health Information Management, School of Health and Rehabilitation Sciences, University of Pittsburgh, Pittsburgh, PA, United States.
JMIR Formative Research
|October 1, 2024
Summary
Generative artificial intelligence (AI) models like ChatGPT show potential in clinical assessments. Prompt engineering improved their ability to administer and interpret the Sour Seven Questionnaire for delirium detection.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing
- Clinical Assessment Tools
Background:
- Generative AI and large language models (LLMs) offer potential for medical education and clinical decision-making.
- ChatGPT, a general-purpose AI, can perform tasks like differential diagnosis without specific training.
- The application of ChatGPT in specialized, context-specific assessment workflows, including scoring and interpretation, requires further study.
Purpose of the Study:
- To evaluate and optimize ChatGPT-3.5 and ChatGPT-4 for administering and interpreting the Sour Seven Questionnaire.
- To train AI models using prompt engineering to accurately apply the Sour Seven Questionnaire to clinical vignettes.
- To assess AI performance against human experts in delirium symptom identification and scoring, refining interpretation accuracy.
Main Methods:
- Prompt engineering was utilized to train ChatGPT-3.5 and ChatGPT-4 on the Sour Seven Questionnaire.
- Specific, structured prompts guided AI models in understanding and applying assessment criteria to clinical vignettes.
- Prompts were designed to ensure standardized response formatting consistent with clinical documentation.
Main Results:
- Both ChatGPT models showed proficiency in applying the Sour Seven Questionnaire, with initial inconsistencies improving over time.
- Iterative prompt engineering enhanced AI performance in detecting delirium symptoms and assigning scores.
- Optimizations included refining scoring methodology, mandating tabular response formats, and ensuring adherence to recommended actions.
Conclusions:
- Preliminary findings support the utility of AI models like ChatGPT for administering standardized clinical assessments.
- Context-specific training and prompt engineering are crucial for maximizing AI potential in healthcare.
- Further research is needed to validate these findings in real-world settings and assess generalizability.
Keywords:
ChatGPTSour Seven Questionnairecaregiver educationclinical vignettesdelirium detectiongenerative AIgenerative artificial intelligencelarge language modelsmedical educationprompt engineering
