Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The Basic Science of Large Language Models in Orthopaedic Surgery
Alan B C Dang1, Alexis B C Dang
1From the Department of Orthopaedic Surgery, University of California San Francisco, San Francisco, CA (Alan B. C. Dang and Alexis B. C. Dang), and the Orthopaedic Section, Surgical Service, San Francisco VA Health Center, San Francisco, CA (Alan B. C. Dang and Alexis B. C. Dang).
Abstract:
Orthopaedic surgeons routinely consult search engines, journals, and curated websites to stay current on orthopaedic knowledge. The emergence of large language models, such as OpenAI ChatGPT and Google MedGemma, is changing the way we search for information and how residents learn. Although many orthopaedic surgeons are users of artificial intelligence (AI), most are uncertain about how these tools actually work and why they sometimes give impressively accurate explanations alongside glaring factual errors and fabricated citations. This review provides an overview of the underlying preclinical studies behind large language models at the level of detail needed to empower orthopaedic surgeons with the knowledge needed to critically evaluate AI outputs, design future research projects, and effectively incorporate AI tools into clinical practice and resident education. Through clinical examples including a Schatzker VI tibial plateau fracture and an L4 pedicle screw sizing question, we illustrate two distinct classes of AI failure-retrieval failures and reasoning failures-and demonstrate how understanding the preclinical studies behind these errors equips surgeons to evaluate any AI tool regardless of where or how it runs.