Related Experiment Video
Updated: Jul 10, 2026

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Evaluating DeepSeek-R1 for gynecological oncology disease consultation in telemedicine: A comparative study with
Haojie Cai1, Yaqian Zhao1, Yongsong Wu1
1Department of Gynecology, Shanghai Key Laboratory of Maternal Fetal Medicine, Shanghai Institute of Maternal-Fetal Medicine and Gynecologic Oncology, Shanghai First Maternity and Infant Hospital, School of Medicine, Tongji University, Shanghai, China.
Objective:
To evaluate the performance and potential of the DeepSeek-R1 in telemedicine consultations for gynecological oncology diseases, comparing its responses with those of human doctors on prominent Chinese online medical platforms.
Methods:
A total of 600 online consultation cases covering four gynecological oncology diseases were collected from "Ding Xiang Doctor" and "Good Doctor Online." After excluding unsuitable cases, 82 were selected. DeepSeek-R1 generated responses based on patients' questions and information, which were anonymized and evaluated alongside human doctors' replies by three professional gynecologists. Seven dimensions were assessed: medical accuracy, clinical applicability, communication effectiveness, safety and compliance, popular science translatability, humanistic care, and overall satisfaction. Statistical analysis was performed using non-parametric tests.
Results:
DeepSeek-R1 significantly outperformed human doctors across all seven evaluation dimensions (p < 0.0001). Among the seven evaluated dimensions, it scored highest in humanistic care, while human doctors scored highest in medical accuracy. Both groups achieved their lowest scores in popular science translatability. DeepSeek-R1's responses were more comprehensive and logically structured but tended to be lengthy, which could increase the cognitive load on patients.
Conclusions:
DeepSeek-R1 demonstrates strong potential in remote gynecological oncology telemedicine, outperforming human doctors in accuracy, applicability, communication, safety, and humanistic care. However, its responses are often lengthy, potentially increasing patient cognitive load, and both DeepSeek-R1 and human doctors show limitations in effectively translating medical knowledge for public understanding. Future work should focus on optimizing the conciseness of LLM responses and enhancing patient-centered communication to improve telemedicine quality and accessibility.