Related Experiment Video
Updated: Oct 3, 2026

A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
Artificial Intelligence-Assisted Current Procedural Terminology Coding in Pediatric Orthopaedic Surgery: Promise or
Krishi S Manem1, Gloria R Gogola2, Rohini M Vanodia2
1McGovern Medical School, The University of Texas Health Science Center at Houston, Houston, TX, USA.
Background:
Accurate assignment of Current Procedural Terminology (CPT) codes is essential for reimbursement, compliance, and quality reporting in pediatric orthopaedic surgery. Manual coding remains resource-intensive and prone to error and variability. Large language models (LLMs) represent a potential solution for automating CPT assignment directly from operative documentation. This study evaluated the performance of ChatGPT-4o and ChatGPT-5 in generating accurate CPT codes from deidentified pediatric orthopaedic operative notes.
Methods:
A retrospective observational study was performed using 20 deidentified pediatric orthopaedic operative notes collected between January and April 2025. Notes were drawn from cases performed by 7 attending pediatric orthopaedic surgeons, representing a diverse academic case mix to expose the models to variability in operative dictation style and documentation standards. Each case was categorized as single-code (1 CPT) or multicode (≥2 CPT). Each operative note was entered into ChatGPT-4o and ChatGPT-5 using a standardized zero-shot prompt without priming or fine-tuning, and model-generated CPT codes were recorded. CPT codes assigned by the institutional billing department served as the gold standard reference. Model outputs were adjudicated as true positives, false positives, or false negatives. Performance metrics included F1 score, precision, recall, exact-match accuracy, and rates of overcoding and undercoding.
Results:
ChatGPT-4o achieved an overall F1 score of 0.48 (precision: 0.45, recall: 0.56), with higher accuracy in single-code cases (F1 score: 0.53) than in multicode cases (F1 score: 0.41). Overcoding occurred in 70% of cases and undercoding occurred in 60%, and exact-match accuracy was 30%. ChatGPT-5 outperformed ChatGPT-4o across all measures, achieving an overall F1 score of 0.69 (precision: 0.67, recall: 0.74), with higher accuracy in single-code cases (F1 score: 0.76) than in multicode cases (F1 score: 0.61). Overcoding occurred in 25% of cases and undercoding occurred in 30%, and exact-match accuracy improved to 60%.
Conclusions:
ChatGPT-5 demonstrated markedly improved performance over ChatGPT-4o in CPT code prediction for pediatric orthopaedic operative notes, particularly in single-code procedures. While accuracy gains were substantial, error rates remain above acceptable thresholds for unsupervised deployment. These findings suggest that LLMs may serve as augmentative tools to streamline coding workflows and reduce administrative burden, with potential earlier applicability in single-code pediatric orthopaedic cases. Larger-scale validation and integration studies are warranted to support broader clinical implementation.
Key Concepts:
(1)ChatGPT-5 substantially outperformed ChatGPT-4o in Current Procedural Terminology (CPT) code prediction from pediatric orthopaedic operative notes, achieving an F1 score of 0.69 versus 0.48 and an exact-match accuracy of 60% versus 30%.(2)Both models performed better on single-code procedures than on multicode procedures, with ChatGPT-5 achieving an F1 score of 0.76 in single-code cases compared to 0.61 in multicode cases.(3)Overcoding was the dominant error pattern for ChatGPT-4o (70% of cases), while ChatGPT-5 substantially reduced both overcoding (25%) and undercoding (30%), suggesting improved contextual bundling recognition.(4)Current error rates remain above acceptable thresholds for unsupervised clinical deployment; however, supervised integration of large language models (LLMs) as a first-pass coding assistant may reduce administrative burden and variability in routine pediatric orthopaedic cases.(5)Larger multi-institutional validation studies with domain-specific fine-tuning and electronic health record integration will be necessary before LLM-assisted CPT coding can be responsibly implemented in clinical billing workflows.
Level Of Evidence:
III, Retrospective Comparative Study.