Related Experiment Video
Updated: Sep 2, 2026

Robotic Left Hepatectomy using Indocyanine Green Fluorescence Imaging for an Intrahepatic Complex Biliary Cyst
Published on: June 24, 2022
Objective technical skills assessment in laparoscopic cholecystectomy: a systematic review of manual, kinematic, and
Elizabeth M A Mainwaring1, Nadia Guidozzi2, Jonnie P James3
1Surgical Intervention Trials Unit, Nuffield Department of Surgical Sciences, University of Oxford, Oxford, UK. lily.mainwaring@msd.ox.ac.uk.
Background:
Technical errors during surgery are a major contributor to preventable adverse outcomes, driving demand for objective assessment tools. Laparoscopic cholecystectomy (LC) is a commonly performed procedure with a clearly defined set of procedure-related complications, making it an ideal candidate for targeted assessment.
Objective:
To evaluate manual, kinematic, and AI-based technical skills assessment tools for LC and compare validity evidence and methodological quality.
Methods:
A PRISMA 2020-compliant systematic review (PROSPERO CRD420251125937) searched MEDLINE, Embase, Web of Science, and Cochrane Library from inception to 12 August 2025. Studies evaluating objective LC skill assessment tools were included. Two reviewers independently performed screening. Data extraction and study appraisal were performed with independent verification by a second reviewer. Validity was assessed using Messick's framework, methodological quality using MERSQI, and risk of bias using COSMIN for manual and kinematic studies and QUADAS-2 for AI studies. Results were synthesised narratively.
Results:
Sixty studies were included (41 manual, 6 kinematic, 13 AI). Most were single centre and retrospective. Manual tools demonstrated the largest body of validity evidence across independent cohorts, though reliability was variable and outcome associations were lacking. Kinematic systems quantified motion but had limited validation. AI systems showed strong internal performance and early real-time use, particularly for critical view of safety detection, but lacked external validation and clinical impact evidence.
Limitations:
Heterogeneity precluded meta-analysis and limited assessment of reporting bias and certainty.
Conclusions:
No tool demonstrated sufficient validity to represent a gold standard. Manual tools are most mature but lack scalability, while AI systems show promise but require robust validation and evidence of clinical impact.