Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance
Oriol Pujol1,2, Robert Ferrer3, Alex Coelho4
1Knee Surgery Unit, Orthopaedic Surgery Department, Vall d'Hebron University Hospital, Universitat Autònoma de Barcelona (Departament de Cirurgia), Barcelona, Spain.
Purpose:
To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA).
Methods:
This is a cross-sectional comparative study. Twenty questions on TKA supported by moderate-strong level of evidence were randomly selected from the WEMA. Three orthopaedic surgeons independently performed a blinded assessment of all LLM-generated responses. An adapted version of the QUEST rating system, a comprehensive framework designed for the objective human assessment of LLM performance across healthcare-related subdomains, was used. Furthermore, the same three evaluators subjectively selected the best-performing LLM response for each question.
Results:
The four LLMs presented statistically significant differences in overall performance based on the QUEST framework (score range 1-5): Gemini 2.5; 4.77 ± 0.06, Claude 4; 4.72 ± 0.07, ChatGPT-5; 4.70 ± 0.08 and GROK 4; 4.62 ± 0.09 (p < 0.001). Gemini 2.5 achieved the highest scores in the Accuracy (4.58 ± 0.70), Comprehensiveness (4.87 ± 0.34) and Trust (4.52 ± 0.62) dimensions. However, Claude 4 obtained the highest score for the Currency (4.23 ± 0.67) dimension. When assessors subjectively selected the superior answer for each question, Claude 4 was chosen most frequently, in 46.7% of cases, followed by ChatGPT-5 in 25.4%, Gemini 2.5 in 22.9% and GROK 4 in 7.5% of cases.
Conclusions:
Four general-purpose LLMs (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) demonstrated good overall performance when addressing specialised clinical questions related to TKA. No single model consistently outperformed the others across all evaluated domains.
Level Of Evidence:
Level V.