Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Introduction to Language of Pathophysiology ll01:17

Introduction to Language of Pathophysiology ll

This lesson explores key terms that describe how diseases progress, their outcomes, and their distribution in populations.Diagnostic tests identify diseases and monitor treatment. These include blood and urine tests, biopsies, imaging (X-ray, MRI), and detection of infectious agents.Remission is a reduction or disappearance of symptoms.Exacerbation refers to the worsening of symptoms, such as increased wheezing during an asthma attack.A precipitating factor triggers an acute episode, while a...
Introduction to Language of Pathophysiology l01:25

Introduction to Language of Pathophysiology l

Pathophysiology investigates how biological mechanisms—typically starting at the cellular level—disrupt normal bodily functions. It bridges anatomy and physiology to explain the progression of disease. With this foundation, it is important to understand the following key terms used to describe disease processes: Diagnosis:The process of identifying a disease using clinical evaluation, including signs (objective evidence like rashes), symptoms (subjective experiences like pain), laboratory test...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Validation of a novel osteoarticular fracture fragment decontamination-preservation system for delayed articular fracture reconstruction for limb preservation in a preclinical canine model pilot study.

Injury·2026
Same author

Detection of Electrophysiologic Conduction After Polyethylene Glycol-Fusion Repair of Peripheral Nerve Injuries in a Porcine Model.

Military medicine·2026
Same author

CORR Insights®: What Is the Interversion Reliability and Agreement Between the Decision Tree Patient-rated Wrist Evaluation and the Full-length Version?

Clinical orthopaedics and related research·2026
Same author

The Role of the Gut Microbiota in Functional Recovery after Peripheral Nerve Injury: A Narrative Review.

Orthopedic reviews·2026
Same author

Measuring What Matters: Patient Perspectives on Success and Recovery After Nerve Injury.

The Journal of hand surgery·2026
Same author

Comparing Conservative and Surgical Treatment of Acute Nondisplaced Fractures of the Scaphoid Based on Fracture Location: A Systematic Review and Meta-Analysis.

The Journal of the American Academy of Orthopaedic Surgeons·2026

Related Experiment Video

Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Assessing Large Language Models for Clinical Coding in Hand Surgery: Effect of Note Authorship, Prompt Design, and

Avery M Schroeder1, Carlye B Goldenberg, Mubinah I Khaleel

  • 1From the School of Medicine, University of Missouri-Columbia (Schroeder, Goldenberg, Khaleel), Columbia, MO, Department of Orthopaedic Surgery, (Schroeder, Khaleel, Nuelle, London) University of Missouri, Columbia, MO, Department of Plastic Surgery (Kirby), University of Missouri, Columbia, MO, and Division of Plastic and Reconstructive Surgery (Kirby), Washington University, St. Louis, MO.

The Journal of the American Academy of Orthopaedic Surgeons
|May 22, 2026
PubMed
Summary

Large language models (LLMs) show potential for medical coding but require optimization. While accurate for CPT codes, LLMs struggle with ICD-10 codes, particularly laterality, indicating they are not yet ready for independent clinical use.

Related Experiment Videos

Last Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Area of Science:

  • Medical informatics
  • Artificial intelligence in healthcare

Background:

  • Assessing Large Language Models' (LLMs) capability in generating accurate medical codes from clinical documentation.
  • Evaluating the impact of note authorship, prompt design, and diagnosis/procedure type on LLM performance.
  • Hypothesizing LLM accuracy exceeding 80% for hand surgery coding.

Purpose of the Study:

  • To evaluate LLM accuracy in generating ICD-10 diagnosis and CPT procedure codes from clinical notes.
  • To determine factors influencing LLM coding performance, including authorship, prompt strategy, and procedure type.
  • To compare the performance of different LLMs (ChatGPT 3.5, ChatGPT 4.0, Gemini) in medical coding tasks.

Main Methods:

  • De-identified clinical and surgical notes from 90 patients undergoing common hand procedures were analyzed.
  • Four prompt types (zero-shot, one-shot, multishot, chain-of-thought) were used to instruct LLMs.
  • Coding correctness rates were calculated and statistically analyzed using Chi-square tests (P < 0.05).

Main Results:

  • LLM performance was consistent across different note authors and prompt types.
  • ChatGPT 3.5 showed lower accuracy for ICD-10 codes compared to ChatGPT 4.0 and Gemini (P < 0.0001).
  • LLMs achieved higher accuracy for CPT codes (91.5%) than ICD-10 codes (23.9%), with laterality errors being common.

Conclusions:

  • Note content variation did not significantly impact LLM coding performance.
  • Public-facing LLMs need further development for accurate clinical documentation coding.
  • Current LLMs are not suitable for independent use in medical coding without significant optimization.