Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmark analysis of myopia-related issues using large language models: a comparison of ChatGPT-4o and deepseek
Jinglei Yao1, Sun Chen Hsin2, Luxi Li1
1Department of Ophthalmology, Beijing Jingmei Group General Hospital, No.18, Heishandajie, Mentougou District, Beijing, 102300, China.
Objective:
This study evaluated the accuracy and comprehensiveness of responses generated by ChatGPT-4o and DeepSeek regarding commonly asked questions about myopia.
Methods:
Thirty myopia-related questions spanning six clinical domains were submitted to both chatbots. Three medical professionals independently rated each response for accuracy and comprehensiveness. Inter-rater reliability was assessed using Fleiss' Kappa, and Shapiro-Wilk tests were conducted to examine normality in rating distributions. Statistical comparisons were performed using the Chi-square test, with significance set at p < 0.05.
Results:
DeepSeek outperformed ChatGPT-4o in overall accuracy, with significantly more responses rated as "Good" (p < 0.0001). Both models demonstrated high comprehensiveness scores when accuracy was rated "Good," though performance declined in treatment-related queries, particularly regarding commercial products like DIMS lenses. Fleiss' Kappa values indicated poor inter-rater agreement (DeepSeek: [Formula: see text] = 0.106; ChatGPT-4o: [Formula: see text] = - 0.0221), and normality tests showed non-normal score distributions (p < 0.0001 across domains).
Conclusion:
Both ChatGPT-4o and DeepSeek can deliver useful responses to myopia-related questions, though limitations remain in areas requiring up-to-date, region-specific treatment information. DeepSeek's stronger performance suggests that localized LLMs may offer competitive advantages. Ongoing refinement, regular data updates, and domain-specific fine-tuning are essential for improving the reliability of AI chatbots in clinical communication.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022