Related Experiment Video
Updated: Apr 23, 2026

Development of Compendium for Esophageal Squamous Cell Carcinoma
Published on: April 12, 2024
Automated data extraction for systematic reviews using GPT-5.2 and Google Gemini Pro 3: A dual-large language model
Prushoth Vivekanantha1, Harjind Kahlon2, Oluwatoba T Balogun2
1Division of Orthopedic Surgery, Department of Surgery, McMaster University, Hamilton, Ontario, Canada.
Purpose:
To evaluate the accuracy, agreement, and efficiency of a dual-large language model (LLM) approach using Generative Pre-Trained Transformer 5.2 (GPT-5.2) and Google Gemini 3 Pro for automated data extraction in orthopaedic systematic reviews.
Methods:
Eight studies from a previously published systematic review on paediatric revision anterior cruciate ligament reconstruction were used to test extraction accuracy, agreement, and efficiency against a pre-defined gold-standard. Both GPT 5.2 and Gemini 3 Pro were prompted via the OpenAI and Google Application Programming Interface (API). Each study had a total of 48 equally weighted data fields to extract from spanning six domains: study characteristics, participant details, injury characteristics, primary and revision surgery details, and outcomes. Extractions were graded as correct, partially correct, or incorrect in reference to the gold-standard.
Results:
Across all 384 fields, both LLMs produced fully correct outputs in 315 (82%) cases, while at least one model was fully correct in 365 (95.1%). Among the six extraction domains, study characteristics (100%, 32/32), injury characteristics (93.8%, 30/32), and outcomes (91.1%, 102/112) showed the highest percentage of at least one model being correct. The entire extraction task was completed in 27 and 35.8 min by GPT-5.2 and Gemini 3 Pro, respectively, for a total API cost of $3.22USD.
Conclusion:
A parallel-LLM approach using GPT-5.2 and Gemini 3 Pro achieved strong accuracy with a high degree of efficiency for automated data extraction in an orthopaedic systematic review. Most errors were due to omission of minor details in complex domains such as surgical details. At least one model was fully correct in over 95% of fields, supporting the use of a dual-LLM framework as a reliable first-pass tool for human verification.
Level Of Evidence:
Level IV.

