Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Generative AI Large Language Models for Urticaria Management: A Comparative Analysis of DeepSeek-R1 and
Mengyao Yang1, Jingchen Liang1, Luyue Zhang1
1Department of Dermatology, the Second Affiliated Hospital of Xi'an Jiaotong University, Xi'an, China.
Introduction:
Urticaria is a prevalent condition affecting a significant portion of the global population. Both dermatologists and patients require access to up-to-date and accurate information. Traditional search engines often fall short in meeting these needs. Despite the growing reliance on AI for medical inquiries, the accuracy and quality of AI-generated remain understudied. This study aims to evaluate and compare the performance of two widely used AI models, ChatGPT-4o and DeepSeek-R1, in addressing urticaria-related queries.
Methods:
An e-Delphi procedure was employed to generate and refine a set of urticaria-related questions, as well as to develop an evaluation framework for AI-generated responses. ChatGPT-4o and DeepSeek-R1 were then prompted with the finalized questions, and their responses were recorded. A single-blind comparative assessment was conducted among 67 participants (29 dermatologists and 38 non-dermatologists). The responses from both AI models were assessed across simplicity, accuracy, professionalism, clinical feasibility, comprehensibility, and completeness.
Results:
DeepSeek-R1 outperformed ChatGPT-4o in most metrics. Dermatologists rated DeepSeek significantly higher in simplicity (p < 0.001), accuracy (p < 0.001), completeness (p = 0.001), professionalism (p < 0.001), and clinical feasibility (p < 0.001). Non-dermatologists found DeepSeek's responses more concise (p < 0.001) and comprehensible (p < 0.001). Both models showed comparable integration of cutting-edge knowledge (p = 0.06), though DeepSeek exhibited greater output stability, as evidenced by lower standard deviations. When compared with the guidelines, the answers provided by DeepSeek-R1 contained no errors, while ChatGPT-4o made errors in three clinical questions.
Conclusion:
AI-generated answers require rigorous evaluation to ensure their reliability and suitability for medical applications. Based on the current study, DeepSeek-R1 outperforms ChatGPT-4o in addressing urticaria-related queries, demonstrating higher potential for both clinical and patient use.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023