Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reasoning-based LLMs surpass average human performance on medical social skills
Khalid Ibraheem Alohali1, Laura Asaad Almusaeeb2, Abdulaziz Abdulrahman Almubarak2
1College of Medicine, King Saud University, Riyadh, 11461, Saudi Arabia. khalid.i.alohali@gmail.com.
Reasoning-based AI models, like o1, excel at medical social skills questions, outperforming other large language models (LLMs) on licensing exams. These advanced AI tools show promise for enhancing medical education and patient care.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Education Technology
- Natural Language Processing
Background:
- Medical licensing exams heavily feature social skills (communication, ethics, professionalism) crucial for patient care.
- Integration of AI in healthcare raises questions about its ability to handle human-centered scenarios.
- Previous studies show large language models (LLMs) perform well on USMLE social skills questions.
Purpose of the Study:
- To evaluate the performance of five LLMs (GPT-4, GPT-4o, Gemini 1.5 Pro, o1-preview, o1) on USMLE-style social skills questions.
- To assess the impact of reasoning capabilities (chain-of-thought) on LLM performance in social skills scenarios.
- To test LLM consistency under skeptical follow-up prompts.
Main Methods:
- Utilized forty USMLE-style social skills questions from the UWORLD question bank.
- Covered domains including communication, healthcare policy, system-based practice, and medical ethics.
- Administered an "Are you sure?" prompt post-answer to evaluate consistency.
Main Results:
- The reasoning-based model o1 achieved the highest accuracy (97.5%), answering 39 out of 40 questions correctly.
- GPT-4o and Gemini 1.5 Pro tied for second place (87.5%), outperforming GPT-4 (75%) and o1-preview (77.5%).
- All evaluated LLMs surpassed the UWORLD question bank average of 64%.
Conclusions:
- Reasoning-based LLMs, particularly o1, demonstrate significant potential for excelling in complex, socially oriented medical tasks.
- GPT-4o and Gemini 1.5 Pro showed distinct domain strengths, highlighting the need for tailored AI applications.
- The consistent high performance of reasoning models suggests their value in complementing clinical training, medical education, and patient care.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Language and Cognition
Factors Influencing Attraction V: Social Skills
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Stereotype Content Model
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...
