Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reason and Intuition01:37

Reason and Intuition

7.5K
The human brain processes information for decision-making using one of two routes: an intuitive system and a rational system (Epstein, 1994; popularized by Kahneman, 2011 as System 1 and System 2, respectively). The intuitive system is quick, impulsive, and operates with minimal effort, relying on emotions or habits to provide cues for what to do next, while the rational system is logical, analytical, deliberate, and methodical. Research in neuropsychology suggests that the...
7.5K
Reliability and Validity01:29

Reliability and Validity

14.1K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.1K
Reasoning01:30

Reasoning

440
Reasoning is the action of thinking about something in a logical, sensible way. It is integral to problem-solving, decision-making, and critical thinking. Reasoning can be inductive or deductive. Reasoning involves transforming information into conclusions, which is essential for problem-solving, decision-making, and critical thinking.
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
440
Deductive Reasoning01:16

Deductive Reasoning

69.1K
Deductive reasoning, or deduction, is the type of logic used in hypothesis-based science. In deductive reasoning, the pattern of thinking moves in the opposite direction as compared to inductive reasoning, which means that it uses a general principle or law to predict specific results. From those general principles, a scientist can deduce and predict the specific results that would be valid as long as the general principles are valid.
For example, a researcher can deduce specific predictions...
69.1K
Language01:16

Language

919
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
919
Inductive Reasoning00:59

Inductive Reasoning

67.9K
Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
67.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Infection following foot and ankle surgery : a subanalysis of data captured from the UK Foot and Ankle Thromboembolism (FATE) audit.

The bone & joint journal·2026
Same author

Pharmacokinetic Prediction of Repurposed Drugs for PDAC Using Artificial Intelligence.

ACS omega·2026
Same author

Effect of TRIPOD+AI Guidelines on the Reporting Quality of Artificial Intelligence Prediction Models in Orthopaedic Surgery: An 18-Month Bibliometric Study.

Cureus·2025
Same author

High Room-Temperature Magnesium Ion Conductivity in Spinel-Type MgYb<sub>2</sub>Se<sub>4</sub> Solid Electrolyte.

Chemistry of materials : a publication of the American Chemical Society·2025
Same author

UK Foot and Ankle Thromboembolism (UK-FATE).

The bone & joint journal·2024
Same author

Design and control of jumping microrobots with torque reversal latches.

Bioinspiration & biomimetics·2024

Related Experiment Video

Updated: Feb 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K

Accuracy Is Not Enough: Reasoning and Reference Reliability in Orthopaedic Large Language Model (LLM) Applications.

Shashwat Singh1, Pranav Chandrasekhar2

  • 1Trauma and Orthopaedics, The Queen Elizabeth Hospital King's Lynn NHS Foundation Trust, King's Lynn, GBR.

Cureus
|February 9, 2026
PubMed
Summary

Large language models like GPT-5 show high accuracy on orthopaedic exams but often fabricate references. Evaluating reasoning and evidence is crucial, not just correct answers.

Keywords:
artificial intelligence in surgerygpt-5large language models (llms)llmorthopaedic surgery

More Related Videos

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
06:16

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting

Published on: June 6, 2020

4.6K
Examining Bilingual Language Control Using the Stroop Task
05:31

Examining Bilingual Language Control Using the Stroop Task

Published on: February 26, 2020

15.6K

Related Experiment Videos

Last Updated: Feb 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K
Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
06:16

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting

Published on: June 6, 2020

4.6K
Examining Bilingual Language Control Using the Stroop Task
05:31

Examining Bilingual Language Control Using the Stroop Task

Published on: February 26, 2020

15.6K

Area of Science:

  • Artificial Intelligence in Medicine
  • Medical Education Technology
  • Orthopaedic Surgery Training

Background:

  • Large language models (LLMs) demonstrate performance on par with postgraduate trainees in orthopaedics.
  • Clinicians increasingly use LLMs for educational and decision-support purposes.
  • Current LLM evaluations focus on accuracy, neglecting reasoning quality and reference reliability.

Purpose of the Study:

  • To systematically evaluate the relationship between answer accuracy, reasoning quality, and reference reliability of LLMs on a postgraduate orthopaedic examination.
  • To assess GPT-5's performance on the 2024 Orthopaedic In-Training Examination (OITE).

Main Methods:

  • GPT-5 was administered the 2024 OITE (203 questions), providing answers, rationales, and references.
  • Accuracy was compared against American Academy of Orthopaedic Surgeons (AAOS) data.
  • A subsample underwent detailed validation of reasoning and references, comparing GPT-5's reasoning to AAOS explanations.

Main Results:

  • GPT-5 achieved 78.3% accuracy, surpassing the OITE pass threshold and PGY-5 resident scores.
  • Hallucinations occurred in 33% of responses, significantly higher in incorrect (50%) versus correct answers (15.9%).
  • Reasoning for correct answers was high, with 95.5% matching AAOS explanations; however, 33% of all answers cited fabricated or misrepresented references.

Conclusions:

  • GPT-5 demonstrates high accuracy on the OITE but exhibits poor reference reliability.
  • Even accurate answers may depend on flawed or unverifiable sources.
  • Evaluating LLMs in medical education requires assessing reasoning and evidence validation alongside accuracy.