Related Experiment Video
Updated: May 4, 2026

06:18
A Bedside, Single Burr Hole Approach to Multimodality Monitoring in Severe Brain Injury
Published on: March 26, 2019
9.1K
Evaluating ChatGPT and DeepSeek in postdural puncture headache management: a comparative study with international
Jiayi Deng1,2, Xu Qiu1,2, Chengqi Dong1,2
1The Fourth Clinical School of Medicine, Zhejiang Chinese Medical University,Hangzhou First People's Hospital, Hangzhou, China.
BMC Neurology
|July 2, 2025
Summary
AI models like ChatGPT and DeepSeek show clinical validity for post-dural puncture headache (PDPH) information. However, their completeness varies, requiring careful healthcare professional oversight before clinical application.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Medical Information Retrieval
Background:
- Post-dural puncture headache (PDPH) is a frequent complication following dural puncture.
- Existing evidence-based guidance for PDPH prevention, diagnosis, and management is limited.
- The increasing use of AI tools by healthcare professionals necessitates evaluation of their clinical information quality.
Purpose of the Study:
- To assess the accuracy and completeness of AI-generated information on PDPH management.
- To compare responses from ChatGPT-4o, ChatGPT-4o mini, DeepSeek-V3, and DeepSeek-R1 against 2023 consensus guidelines.
Main Methods:
- Evaluation of AI model responses against PDPH consensus guidelines across four dimensions: accuracy, overconclusiveness, supplementary information, and incompleteness.
- Utilized a 5-point Likert scale to assess response accuracy and completeness.
- Compared responses from ChatGPT-4o, ChatGPT-4o mini, DeepSeek-V3, and DeepSeek-R1.
Main Results:
- All four AI models demonstrated 100% accuracy in guideline adherence for PDPH.
- No models provided overly conclusive or unjustified recommendations.
- Response completeness varied, with ChatGPT-4o and DeepSeek-R1 showing higher guideline alignment (70-80%) compared to ChatGPT-4o mini and DeepSeek-V3 (60%).
Conclusions:
- ChatGPT-4o and DeepSeek-R1 exhibit strong clinical validity and guideline alignment for PDPH information.
- Despite high accuracy, AI models provide only partial guideline coverage (60-80% completeness), mandating critical human evaluation.
- Further research is crucial to establish reliable AI support for clinical decision-making in complex conditions like PDPH.

