Related Experiment Video
Updated: Jun 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Custom and Off-the-Shelf Large Language Models Routinely Misinterpret Implant Technique Guides: Too Soon to
Joshua J Woo1, Andrew J Yang1, Yash S Saboo2
1The Warren Alpert Medical School of Brown University, Providence, Rhode Island.
Background:
Large language models (LLMs) are used for clinical information retrieval, yet their performance on highly domain-specific documents such as orthopaedic technique guides or instructions for use (IFU) remains poorly understood. Various financial drivers affecting the orthopaedic medical device industry have generated interest in automated perioperative support during surgery using advanced generative artificial intelligence (AI) techniques leveraging LLMs. We sought to establish whether these complex, manufacturer-specific IFUs for surgical planning and intraoperative execution were clinically amenable to substitution by custom LLM applications.
Methods:
We evaluated 5 LLM-based information retrieval solutions, including 4 custom retrieval-augmented generation pipelines and ChatGPT5, in their ability to extract clinically relevant information from 3 distal femoral replacement IFUs. Two fellowship-trained orthopaedic surgeons curated 28 questions spanning literal, enumerative, and reasoned query types. Answers were scored against the expert-generated ground truth using a three-tier rubric (incorrect, partially correct, fully correct).
Results:
All systems demonstrated low overall accuracy (<50%). A custom multimodal pipeline achieved the highest overall score (44.6%), outperforming commercial systems such as ChatGPT (29.2%). Performance varied by document and question type: literal queries were most accurately answered (up to 53.0%), while reasoned questions yielded the lowest scores across all systems (as low as 15.3%).
Conclusions:
Current LLM-based retrieval systems, including commercially available tools, are unreliable for extracting complex procedural information from orthopaedic implant protocols. Whether IFUs require further clarity for clinical queries or open source LLMs require enhanced image processing, medical device representatives are far from being replaced by generative AI techniques given their poor performance in safe integration with surgical workflows. Improving LLMs with enhanced image processing and domain-specific training will be necessary before considering medical device representatives' substitution.
Level Of Evidence:
Prognostic, Level IV. See Instructions for Authors for a complete description of levels of evidence.
