Related Experiment Video
Updated: Sep 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Models in Ptosis-Related inquiries: A Cross-Lingual Study
Ling-Han Niu1, Li Wei2, Bixuan Qin3
1Beijing Tongren Eye Center, and Beijing Ophthalmology Visual Science Key Lab, Beijing Tongren Hospital, Capital Medical University, Beijing, People's Republic of China.
Large language models (LLMs) like GPT-4o show strong performance for ptosis questions, improving patient education and clinician support. Qwen2.5 excels in Chinese clarity, highlighting the need for validation before clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology Research
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly explored for healthcare applications.
- Ptosis inquiries require accurate, clear, and empathetic responses for both patients and clinicians.
- Cross-lingual capabilities are crucial for global healthcare accessibility.
Purpose of the Study:
- To evaluate the performance of GPT-4, GPT-4o, Qwen2, and Qwen2.5 in addressing ptosis-related questions.
- To assess the cross-lingual applicability and patient-centric assessment capabilities of these LLMs.
- To compare LLM performance on both patient-focused and clinician-focused inquiries.
Main Methods:
- Collected 11 patient-centric and 50 clinician-centric ptosis questions.
- Evaluated LLM responses using criteria like accuracy, sufficiency, clarity, depth, helpfulness, and empathy.
- Conducted clinical assessments with 30 patients and 8 oculoplastic surgeons using a 5-point Likert scale.
Main Results:
- GPT-4o outperformed Qwen2.5 in overall performance and completeness for clinician questions.
- GPT-4o scored higher in helpfulness for patient questions, with no significant differences in clarity or empathy.
- Qwen2.5 demonstrated superior clarity in Chinese compared to English.
Conclusions:
- LLMs, especially GPT-4o, show significant potential for ptosis-related inquiries, offering valuable insights.
- Qwen2.5 shows promise for Chinese language support, indicating multilingual strengths.
- Further validation, domain-specific training, and cultural adaptation are necessary before clinical deployment of LLMs.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018