Related Experiment Videos
Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in
Marco Di Maio1,2, Giorgio Stopper3, Vincenzo Di Matteo1,2
1Department of Biomedical Sciences, Humanitas University, Via Rita Levi Montalcini 4, Pieve Emanuele, 20072 Milan, Italy.
Background:
The ASA Physical Status (ASA-PS) classification and the Charlson Comorbidity Index (CCI) are common pre-operative scoring tools. Language models could automate structured pre-operative scoring, but direct comparisons require paired inference because all models are evaluated on the same patients.
Methods:
In this retrospective single-center concordance analysis, 101 consecutive adult orthopedic patients were independently rated by two clinicians; the rounded mean for ASA-PS and arithmetic mean for CCI formed a clinician-derived composite reference. The cohort contained no ASA-PS IV-V patients. Six model configurations received identical prompts. Agreement was assessed using quadratic weighted kappa, ICC(2,1), exact and adjacent agreement, MAD, RMSE, and Bland-Altman limits. Post hoc between-model comparisons used 10,000 patient-level paired bootstrap replicates with Benjamini-Hochberg correction.
Results:
Inter-clinician weighted kappa was 0.713 for ASA-PS and 0.914 for CCI. GPT-5.2 reached kappa 0.884 for ASA-PS and 0.970 for CCI. In paired analyses, GPT-5.2 had significantly higher quadratic weighted kappa than every other tested model for both outcomes and significantly higher ICC for CCI after multiplicity correction. Phi4 and deepseek-r1-70B were not significantly different from inter-clinician agreement for CCI kappa or ICC; equivalence was not tested.
Conclusions:
Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction. Locally deployable phi4 and deepseek-r1-70B showed CCI agreement estimates that were not statistically distinguishable from inter-clinician agreement, although equivalence was not tested. These findings are limited to the evaluated models and study cohort.