Evaluating ChatGPT's Role in Supporting Military Readiness Assessment
Ashley M Tyburski1, Ensign Sean P Garin1, Col Justin P Fox2,3
1School of Medicine, Uniformed Services University, Bethesda, MD 20814, United States.
Introduction:
Joint Knowledge, Skills, and Abilities (JKSA) scoring of clinical workload data is a robust metric for military readiness assessment. However, the calculation of these scores requires a data analysis skillset that is not widely available. To address this gap, we developed a custom-generative pretrained transformer (GPT) model for JKSA scoring and compared performance to a "gold standard."
Materials And Methods:
To conduct the study, we utilized de-identified, clinical workload data from a single military treatment facility's military-civilian partnership program collected from January to December 2023. First, the data were divided into training, validation, and test datasets. Second, the custom GPT was trained for JKSA calculation on the training set. Then, it was refined using the validation set. Finally, the 3 authors used the GPT to independently calculate JKSA scores, patient encounters, procedure counts, and critical care encounters for the test set. The correlation coefficient (CC) was calculated to quantify the agreement between the GPT's scoring and scoring using traditional techniques within the SAS, version 9.4 environment.
Results:
The overall dataset contained information for 22,811 patient encounters performed by 40 providers in 8 critical wartime specialties. The GPT-calculated diagnostic (CC = 0.92, P < .001) and procedural (CC = 0.76, P < .001) JKSA scores were significantly correlated with those from standard processes. Similarly, GPT-calculated patient encounters (CC = 0.98, P< .001), procedures (CC = 1.00, P < .001), and critical care encounters (CC = 1.00, P < .001) were significantly correlated with SAS calculations. When using the GPT model, we identified key lessons learned for data management, prompt engineering, and cross-checking to facilitate the model's success.
Conclusions:
The custom-GPT model proved an accurate method to calculate JKSA scores from clinical workload data to support military readiness assessment. This project represents a step forward in making the JKSA metric more widely accessible. Further research is required to test performance among new, less homogenous datasets and additional users.


