Related Experiment Video
Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing Large Language Models for Clinical Coding in Hand Surgery: Effect of Note Authorship, Prompt Design, and
Avery M Schroeder1, Carlye B Goldenberg, Mubinah I Khaleel
1From the School of Medicine, University of Missouri-Columbia (Schroeder, Goldenberg, Khaleel), Columbia, MO, Department of Orthopaedic Surgery, (Schroeder, Khaleel, Nuelle, London) University of Missouri, Columbia, MO, Department of Plastic Surgery (Kirby), University of Missouri, Columbia, MO, and Division of Plastic and Reconstructive Surgery (Kirby), Washington University, St. Louis, MO.
Background:
This study sought to assess large language models' (LLM) ability to generate correct ICD-10 and CPT codes using clinical documentation, and to determine whether note authorship, prompt design, or diagnosis/procedure type affect LLM performance. We hypothesized that LLMs can code hand surgery clinic and surgical notes with greater than 80% accuracy.
Methods:
Ninety patients evenly distributed across three orthopaedic hand surgeons and procedure types (cubital tunnel, carpal tunnel, and trigger finger release) were identified. Clinic and surgical notes were deidentified, and correct ICD-10 diagnosis and CPT procedure codes were recorded. "Zero-shot," "one-shot," "multishot," and "chain-of-thought" prompts instructed LLMs to assign ICD-10 codes and CPT codes based on note content. Each prompt was posed to Chat GPT 3.5, Chat GPT 4.0, and Gemini. Rates of coding correctness were calculated across attendings, diagnosis/procedure, prompt type, and LLM. Chi-square analysis determined statistical significance for these comparisons ( P < 0.05).
Results:
No differences in LLM coding performance were observed between note authors ( P = 0.09 ICD-10, P = 0.48 CPT) or prompt types ( P = 0.27 ICD-10, P = 0.62 CPT). Chat GPT 3.5 provided less accurate ICD-10 codes than Chat GPT 4.0 or Gemini ( P < 0.0001). All LLMs better predicted CPT codes (91.5% correct) than ICD-10 codes (23.9% correct). The most common error was incorrect or omitted ICD-10 laterality. Prompts updated to emphasize ICD-10 laterality demonstrated improved accuracy (40%).
Discussion:
Variation in note content and writing style did not markedly affect LLM performance. Public-facing LLMs require additional optimization to interpret clinical documentation for coding purposes and are not ready for independent use.
Related Concept Videos
Introduction to Language of Pathophysiology ll
Introduction to Language of Pathophysiology l