Evaluating data extraction error by a large language model from randomised controlled trials: a large-scale empirical

Shiqi Fan1, Ming Chen2, Suhail A Doi3

  • 1Proof of Concept Center, Shanghai Eastern Hepatobiliary Surgery Hospital, Shanghai, China.

Summary

Large language models (LLMs) like Claude 3.5 Sonnet show low data extraction error rates from randomized controlled trials (RCTs). However, careful verification of LLM outputs is crucial for evidence synthesis applications.