Related Experiment Video
Updated: Apr 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Input structure-driven instability and convergence in large language model clinical reasoning: a formative study
Vinson James1, Catherine Caronia2, Rajesh Savargaonkar2
1Department of Pediatrics, Good Samaritan University Hospital, West Islip, NY, 11795, USA. dr.vinsonjames@gmail.com.
Large language models (LLMs) show unstable clinical reasoning when given many questions at once. Structuring questions in batches significantly improves LLM performance consistency and reliability in medical education.
Area of Science:
- Artificial Intelligence in Medical Education
- Clinical Reasoning Assessment
- Large Language Model (LLM) Performance
Background:
- Large language models (LLMs) are increasingly used in medical education and clinical settings.
- Prior research focused on LLM accuracy on exams, but not on the stability of clinical reasoning with varied input structures.
- Input structure is critical for safe educational and clinical deployment of LLMs.
Purpose of the Study:
- To investigate how question delivery structure affects performance stability, inter-model variability, and reproducibility of LLM clinical reasoning.
- To evaluate contemporary LLMs using pediatric residency-level multiple-choice questions (MCQs).
Main Methods:
- Generated 77 validated pediatric USMLE Step 2/3-style MCQs emphasizing diagnostic, management, and ethical reasoning.
- Evaluated six publicly available LLMs (October-December 2025 versions) under two conditions: simultaneous presentation vs. sequential delivery in batches of ten.
- Compared accuracy and inter-model variability using paired t-tests and one-way ANOVA.
Main Results:
- Simultaneous question presentation showed wide accuracy variation (38%-90%) and poor reproducibility across models.
- Sequential batch delivery improved performance convergence (83%-88%) with no significant inter-model differences.
- Batch delivery substantially reduced performance dispersion and instability across all evaluated LLMs.
Conclusions:
- LLM clinical reasoning is highly sensitive to input structure; prompt structure is key to reliable behavior.
- Structured batch delivery minimizes contextual load, improving reproducibility and reducing inter-model variability.
- Consider prompt structure in designing AI-supported medical education and assessment systems for reliable LLM use.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Introduction to Language of Pathophysiology ll
Critical Thinking II
Introduction to Language of Pathophysiology l
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Mathematical Modeling: Problem Solving
Patient-centered Care