Related Experiment Video
Updated: Jun 2, 2026

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
Published on: January 29, 2020
Retrieval-augmented generation enhances large language model performance on the Japanese orthopedic board examination
Juntaro Maruyama1, Satoshi Maki2, Takeo Furuya1
1Department of Orthopedic Surgery, Graduate School of Medicine, Chiba University, Japan.
Introduction:
Large language models (LLMs) have shown potential in medical applications. However, their effectiveness in specialized medical domains remains underexplored. The integration of Retrieval-Augmented Generation (RAG) has been proposed to improve these models by reducing hallucinations and enhancing domain-specific information access. Through this evaluation, we aim to assess whether RAG can effectively bridge the gap between LLMs' current capabilities and the accuracy needed for medical use by examining GPT-3.5 Turbo, GPT-4o, and o1-preview on the 2024 Japanese Orthopedic Specialist Examination.
Methods:
A specialized database was created using the "Standard Textbook of Orthopedics", and GPT-3.5 Turbo, GPT-4o, and o1-preview were evaluated with and without RAG. Models were tested on text-based and image-based questions exactly as presented in Japanese. An error analysis was conducted to identify key performance factors.
Results:
GPT-3.5 Turbo showed no substantial improvement with RAG, with its overall accuracy remaining at 28 %, compared to its baseline of 29 % without RAG. GPT-4o rose from 62 % to 72 %, while o1-preview increased from 67 % to 84 %. Error analysis indicated that GPT-3.5 Turbo primarily failed to apply retrieved data, whereas GPT-4o and o1-preview made errors when the database lacked relevant information or when dealing with image-based questions.
Conclusions:
The integration of RAG significantly boosted performance for GPT-4o and especially o1-preview. While both models surpassed the passing threshold, o1-preview demonstrated a level of proficiency relevant to clinical practice. However, RAG did not improve performance on GPT-3.5 Turbo because it lacks effective reasoning abilities.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Genetic Lingo
Improving Translational Accuracy
Improving Translational Accuracy
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...