Can a Large Language Model Interpret Data in the Electronic Health Record to Infer Minimum Clinically Important
Abdul K Zalikha1, Thomas S Hong1, Easton A Small1
1Department of Orthopedic Surgery, Stanford University and VA Palo Alto Health Care System, Palo Alto, California.
Background:
Obtaining total knee arthroplasty patient-reported outcomes for quality assessment is costly and difficult. We asked whether a large language model (LLM) could interpret electronic health record notes to differentiate patients attaining a 1-year minimum clinically important difference (MCID) for the Knee Osteoarthritis Outcome Score-Joint Replacement (KOOS-JR) from those who did not. We also investigated whether sufficient information to infer MCID achievement exists in the chart by having a blinded orthopaedic surgeon make the same determination.
Methods:
In this retrospective case-control study, we selected 40 total knee arthroplasty patients who achieved 1-year KOOS-JR MCID and 40 who did not. Orthopaedic, emergency medicine, and primary care notes from zero to six months preoperatively and nine to 15 months postoperatively were deidentified. ChatGPT 3.5 (ChatGPT) interpreted these notes to determine whether the patient improved after surgery. A blinded orthopaedic surgeon classified these patients using all chart information. The sensitivity, specificity, and accuracy of ChatGPT and the surgeon's responses were calculated.
Results:
ChatGPT classified 78 of 80 cases with 97% sensitivity, but only 33% specificity. The surgeon's assessment had 90% sensitivity and 63% specificity. Given the equal distribution of patients meeting or not meeting MCID, Chat GPT's accuracy was 65%. The surgeon's was 76%.
Conclusions:
ChatGPT's assessment of KOOS-JR MCID attainment had 97% sensitivity, but only 33% specificity. False positives were commonly due to the LLM not having access to, or not properly interpreting, signs of problems in the chart. This was an initial evaluation of the current ability of a general-purpose LLM to evaluate patient outcomes based on information in chart notes. An orthopaedic surgeon's assessment of the full chart suggests an opportunity to improve on this baseline performance, possibly enabling quality monitoring and identification of best practices across a large health care system. Additional work is needed to optimize model performance and confirm the utility of this approach.
Related Concept Videos
Improving Translational Accuracy
Methods of Documentation VII: EMR
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Knee Joint
A total of seven ligaments support the knee joint. The patellar ligament, which is also attached to the quadriceps femoris...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Purpose of Health Records II


