Related Experiment Video
Updated: Jul 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Therapy Recommendations Across 3 Clinical Specialties: Comparative Study.
Theresa Isabelle Wilhelm1,2, Jonas Roos3, Robert Kaczmarczyk4,5
1Eye Center, Medical Center, Faculty of Medicine, University of Freiburg, Freiburg, Germany.
Large language models (LLMs) show promise in generating medical information but require careful evaluation for accuracy and safety. Claude-instant-v1.0 performed best, while GPT-3.5-Turbo was safest, highlighting the need for ongoing AI assessment in healthcare.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly used for medical information generation.
- Rigorous assessment of AI-generated medical content quality, accuracy, and safety is crucial across specialties.
- The rapid advancement of AI necessitates evaluating its role in healthcare.
Purpose of the Study:
- To evaluate the medical content generation performance of four prominent LLMs: Claude-instant-v1.0, GPT-3.5-Turbo, Command-xlarge-nightly, and Bloomz.
- To assess AI-generated therapeutic recommendations in ophthalmology, orthopedics, and dermatology.
- To compare physician evaluations with automated GPT-4 assessments for LLM-generated medical content.
Main Methods:
- Physician evaluation of AI-generated therapeutic recommendations for 60 diseases using mDISCERN scores, correctness, and harmfulness.
- Statistical analysis (ANOVA, t-tests) to compare model and specialty performance.
- Automated evaluation using GPT-4, compared to physician assessments via Pearson correlation.
Main Results:
- Claude-instant-v1.0 achieved the highest mDISCERN score; Bloomz had the lowest.
- Significant differences in content quality and safety were observed across LLMs and specialties.
- GPT-3.5-Turbo demonstrated the lowest harmfulness rating; common errors included diagnostic confusion and omitted treatments.
- GPT-4 evaluations showed substantial alignment with physician assessments.
Conclusions:
- LLMs demonstrate capability in generating medical content, but quality and safety require refinement.
- Regular, methodical assessments and oversight are essential for trustworthy AI medical advice.
- GPT-4 offers a scalable method for automated, domain-agnostic evaluation of AI-generated medical content.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Language and Cognition
Treatment Strategies for Psychological Disorders
Psychological therapies focus on modifying emotions, thoughts, and behaviors through talking, interpreting, listening, rewarding, challenging, and modeling. Clinical psychologists, counselors, and social workers commonly practice psychotherapy. Clinical...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Typical Model Studies
Components of Language