Related Experiment Video
Updated: Jun 27, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Comparing manual vs. automated machine learning and deep learning models for predicting one-year mortality in elderly
Adi Shuchami1, Maxim Glebov2,3, Maksim Katsin2
1Department of Mathematics, Ariel University, Ariel, Israel.
Background:
Hip fractures are associated with significant mortality, especially among elderly patients. Accurate prediction of mortality risk is crucial for optimising perioperative care and resource allocation. Recent advances in machine learning (ML) and deep learning (DL) offer promising methods to enhance clinical risk prediction models; however, their clinical implementation often remains limited due to the complexity of these techniques.
Methods:
This retrospective cohort study included 2,604 elderly patients (≥65 years) undergoing urgent hip fracture surgery at Sheba Medical Center, Israel, between January 2017 and November 2023. Multiple ML and DL algorithms were evaluated for predicting one-year all-cause mortality using a comprehensive set of clinical, demographic, perioperative, and laboratory variables. Models were rigorously developed and validated through stratified 5-fold cross-validation, addressing class imbalance with the Synthetic Minority Oversampling Technique (SMOTE). Additionally, an automated ML pipeline, generated using a large language model (LLM) coupled with the Tree-based Pipeline Optimisation Tool (TPOT), was benchmarked against manually optimised models. Model performances were assessed using area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, F1-score, false-positive rate, and true-negative rate, supplemented by permutation importance and SHapley Additive exPlanations (SHAP) for interpretability.
Results:
Among all models evaluated, the manually optimised Extreme Gradient Boosting (XGB) algorithm demonstrated superior predictive performance (AUC = 0.846, accuracy = 0.791, F1-score = 0.667, precision = 0.773, NPV = 0.798). Important predictors identified included baseline serum albumin and urea levels, patient age, intraoperative hypothermia, and the number of chronic diseases. The automated ML model, generated via LLM and TPOT frameworks, showed comparable performance to the XGB model (AUC = 0.844), with a higher recall but slightly lower precision.
Discussion:
ML-based models, particularly the XGB algorithm, significantly enhance predictive accuracy for one-year mortality among elderly hip fracture patients. Crucially, an automated ML framework leveraging large language models provides a practical, clinically accessible alternative, effectively democratising advanced predictive analytics in healthcare settings.