Related Experiment Videos
Predicting VMAT patient-specific quality assurance outcome using a machine learning workflow
Suvankar Das1,2, A Robert Xavier3, B Paul Ravindran4
1Department of Physics, St Joseph University, Chumoukedima, Nagaland, India. Suvankar.das@cihsr.ac.in.
Abstract:
To develop a machine learning (ML)-based workflow for virtual patient-specific quality assurance (PSQA) to predict local gamma passing rates (GPRs) and QA pass/fail outcomes for Volumetric Modulated Arc Therapy (VMAT) plans. Treatment plan files from 156 double-arc VMAT cases were used, resulting in 312 beam-specific datasets. A total of 48 plan complexity metrics (PCMs), derived from DICOM RT plan files based on the UCoMX framework, were calculated for each arc. Feature selection was performed using a blended ensemble of Random Forest (RF), Gradient Boosting Decision Tree (GBDT), eXtreme Gradient Boost (XGBoost), and linear Support Vector Machine (SVM) models, yielding 15 key predictive features. Regression models using RF, GBDT, XGBoost, SVR, and an averaging ensemble were trained to predict local GPR values. Patient-level grouped train-test splitting and grouped 5-fold cross-validation were employed to prevent data leakage. Classification models were developed using a 90% GPR threshold to predict QA pass/fail outcomes. Model performance was evaluated using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), ROC-AUC, sensitivity, specificity, Positive Predictive Value (PPV), and Negative Predictive Value (NPV). Among the regression models, GBDT achieved the best overall performance with an MAE of 1.85% points and an RMSE of 2.35% points, followed closely by the averaging ensemble model. Grouped 5-fold cross-validation demonstrated stable regression performance across all models. For classification analysis, grouped 5-fold cross-validation ROC-AUC values ranged from 0.721 to 0.772, with SVM achieving the highest mean cross-validation ROC-AUC. On the independent test dataset, SVM achieved the highest ROC-AUC (0.725) and specificity (0.829), while the averaging ensemble model demonstrated balanced classification performance with improved specificity and PPV. An ML-based workflow for predicting VMAT PSQA outcomes from plan complexity metrics derived from DICOM RT plans has been developed. The results demonstrate the feasibility of predicting QA outcomes with moderate accuracy and provide a framework for further development of ML-assisted PSQA. Future studies incorporating additional plan, delivery, and machine-related features, together with larger multi-institutional datasets, may further improve predictive performance and clinical applicability.