Related Experiment Video
Updated: Jun 9, 2025

The Multiple Sclerosis Performance Test MSPT: An iPad-Based Disability Assessment Tool
Published on: June 30, 2014
Performance Assessment of GPT 4.0 on the Japanese Medical Licensing Examination
Hong-Lin Wang1,2, Hong Zhou1,2, Jia-Yao Zhang1,2
1Department of Orthopedics Surgery, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, 430022, China.
Objective:
To evaluate the accuracy and parsing ability of GPT 4.0 for Japanese medical practitioner qualification examinations in a multidimensional way to investigate its response accuracy and comprehensiveness to medical knowledge.
Methods:
We evaluated the performance of the GPT 4.0 on Japanese Medical Licensing Examination (JMLE) questions (2021-2023). Questions are categorized by difficulty and type, with distinctions between general and clinical parts, as well as between single-choice (MCQ1) and multiple-choice (MCQ2) questions. Difficulty levels were determined on the basis of correct rates provided by the JMLE Preparatory School. The accuracy and quality of the GPT 4.0 responses were analyzed via an improved Global Qualily Scale (GQS) scores, considering both the chosen options and the accompanying analysis. Descriptive statistics and Pearson Chi-square tests were used to examine performance across exam years, question difficulty, type, and choice. GPT 4.0 ability was evaluated via the GQS, with comparisons made via the Mann-Whitney U or Kruskal-Wallis test.
Results:
The correct response rate and parsing ability of the GPT4.0 to the JMLE questions reached the qualification level (80.4%). In terms of the accuracy of the GPT4.0 response to the JMLE, we found significant differences in accuracy across both difficulty levels and option types. According to the GQS scores for the GPT 4.0 responses to all the JMLE questions, the performance of the questionnaire varied according to year and choice type.
Conclusion:
GTP4.0 performs well in providing basic support in medical education and medical research, but it also needs to input a large amount of medical-related data to train its model and improve the accuracy of its medical knowledge output. Further integration of ChatGPT with the medical field could open new opportunities for medicine.
More Related Videos
05:04Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
Published on: August 9, 2024
06:28E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Related Concept Videos
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:
Assessment of the Abdomen II: Percussion
Percussion
Percussion is an essential...
Pulmonary Function Tests
Pulmonary Function Tests are crucial diagnostic tools for assessing respiratory function, particularly in patients with chronic respiratory disorders. They comprehensively evaluate lung volumes, ventilatory function, breathing mechanics, diffusion, and gas exchange. These tests help diagnose pulmonary diseases and play a significant role in monitoring disease progression, evaluating disability, and assessing response to therapy.
PFTs involve using a spirometer, a...
Data Collection III
The principles to begin the physical assessment include conducting a comprehensive or problem-related history in a quiet, well-lit room, emphasizing privacy and comfort for the...
Assessment of the Cardiovascular System III: Palpation
Jugular Venous Pressure (JVP) Measurement
Position the patient at a thirty- to forty-five-degree angle or in a semi-fowler's position. Look for the highest point of pulsation in the internal jugular vein and measure the vertical distance to the angle of Loius or sternal angle. A normal JVP is 3-4 cm above...
Assessment of the Abdomen III: Palpation