Related Experiment Video
Updated: May 23, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for zero-shot procedure extraction in orthopedic surgery: a comparative evaluation.
Ashton Williamson1, Nazgol Tavabi1, Nishita Kalepalli1
1Department of Orthopedics and Sports Medicine, Boston Children's Hospital, Harvard Medical School, Boston, MA, 02115, United States.
Scientific Reports
|May 21, 2026
Summary
Large language models (LLMs) can now extract detailed surgical procedure data from electronic health records, outperforming manual coding. While effective for common procedures, challenges remain for rare cases, requiring further development for expert-level accuracy.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Data Extraction
Background:
- Manual coding of operative notes is inefficient, costly, and inconsistent.
- Electronic health records contain vital surgical care information.
- Large language models (LLMs) offer potential for automated information extraction.
Purpose of the Study:
- Evaluate state-of-the-art LLMs for zero-shot structured information extraction from orthopedic clinical notes.
- Compare LLM performance against human annotations and administrative coding.
- Analyze the impact of model scale, reasoning, and prompt design on extraction accuracy.
Main Methods:
- Tested 14 open-source and proprietary LLMs on 800 orthopedic operative notes.
- Notes were annotated by an orthopedic surgeon and administrator for 74 procedure classes.
- Assessed model outputs against human annotations using macro-F1 scores.
Main Results:
- LLMs consistently outperformed administrator-assigned labels, with macro-F1 scores above 0.6.
- Performance improved by up to 10 points compared to administrative coding.
- Larger models (up to 30 billion parameters) and reasoning capabilities enhanced accuracy, with performance varying by procedure frequency.
Conclusions:
- Modern LLMs show promise for faster, cheaper, and more consistent surgical data curation.
- Full alignment with surgical experts, especially for rare procedures, remains an open challenge.
- General-purpose LLMs can advance automated clinical data curation and surgical informatics.