Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Positron Emission Tomography01:29

Positron Emission Tomography

3.9K
Positron emission tomography (PET) is a medical imaging technique involving radiopharmaceuticals — substances that emit short-lived radiation. Although the first PET scanner was introduced in 1961, it took 15 more years before radiopharmaceuticals were combined with the technique and revolutionized its potential.
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
3.9K
The Availability Heuristic01:08

The Availability Heuristic

5.9K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
5.9K
The Scientific Method02:40

The Scientific Method

59.0K
Research is what makes the difference between facts and opinions. Facts are observable realities, and opinions are personal judgments, conclusions, or attitudes that may or may not be accurate. In the scientific community, facts can be established only using evidence collected through empirical research.
59.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Challenging the Treatment Threshold: Early Faricimab Prevents Vision Loss in nAMD - A Real-World Welsh Experience.

Clinical ophthalmology (Auckland, N.Z.)·2026
Same author

Is reflective writing a valid summative assessment in the era of GenAI? A scoping review.

Medical teacher·2026
Same author

Real-World Outcomes of Switching to Faricimab in Treatment-Experienced and Resistant Neovascular Age-Related Macular Degeneration: A Single-Centre Retrospective Study.

Clinical ophthalmology (Auckland, N.Z.)·2026
Same author

Correction: Grant et al. Low pH, High Stakes: A Narrative Review Exploring the Acid-Sensing GPR65 Pathway as a Novel Approach in Renal Cell Carcinoma. <i>Cancers</i> 2025, <i>17</i>, 3883.

Cancers·2026
Same author

Diabetic Eye Screening: Evaluation and Comparison of the Grading Results Between Community and Hospital Eye Services.

Clinical optometry·2025
Same author

Low pH, High Stakes: A Narrative Review Exploring the Acid-Sensing GPR65 Pathway as a Novel Approach in Renal Cell Carcinoma.

Cancers·2025

Related Experiment Video

Updated: May 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

465

Can ChatGPT-4o Really Pass Medical Science Exams? A Pragmatic Analysis Using Novel Questions.

Philip M Newton1, Christopher J Summers1, Uzman Zaheer1

  • 1Swansea University Medical School, Swansea, Wales, SA2 8PP UK.

Medical Science Educator
|May 12, 2025
PubMed
Summary

ChatGPT-4o demonstrates high performance on medical licensing exams, including novel questions. Its capabilities challenge academic integrity in unproctored settings, necessitating secure exam environments.

Keywords:
Academic integrityAssessment validityCheatingEvidence-based educationMCQsPragmatism

More Related Videos

The Multiple Sclerosis Performance Test MSPT: An iPad-Based Disability Assessment Tool
11:35

The Multiple Sclerosis Performance Test MSPT: An iPad-Based Disability Assessment Tool

Published on: June 30, 2014

57.7K
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
00:08

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

6.9K

Related Experiment Videos

Last Updated: May 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

465
The Multiple Sclerosis Performance Test MSPT: An iPad-Based Disability Assessment Tool
11:35

The Multiple Sclerosis Performance Test MSPT: An iPad-Based Disability Assessment Tool

Published on: June 30, 2014

57.7K
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
00:08

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

6.9K

Area of Science:

  • Artificial Intelligence in Medical Education
  • Natural Language Processing in Healthcare

Background:

  • Concerns exist regarding ChatGPT's potential for academic misconduct in unproctored online medical licensing exams.
  • Previous studies suggested ChatGPT's performance might be inflated by training data and limited by image-based questions.

Purpose of the Study:

  • To evaluate the performance of ChatGPT-4o on existing and novel medical licensing exam questions.
  • To assess the impact of novel question formats and image-based questions on ChatGPT-4o's performance.

Main Methods:

  • ChatGPT-4o was tested on United Kingdom Medical Licensing Exam Applied Knowledge Test and United States Medical Licensing Exam Step 1 questions.
  • Performance was evaluated on existing, rewritten, and entirely novel exam questions, including those with images.

Main Results:

  • ChatGPT-4o achieved high scores: 94% on the UK exam and 89.9% on the US exam.
  • Performance remained strong on novel and rewritten questions, indicating adaptability beyond training data.
  • A slight decrease in performance was observed on image-based questions with text labels.

Conclusions:

  • ChatGPT-4o exhibits advanced capabilities in medical licensing exams, even with novel content.
  • The findings underscore the need for secure testing environments to ensure valid assessment and mitigate academic dishonesty.
  • Continuous improvement in AI models necessitates ongoing reevaluation of assessment strategies in medical education.