Related Experiment Video
Updated: Oct 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis
Jesse P Caron1, Oscar Rodriguez1, Leila Raden2
1Department of General Surgery, AdventHealth Orlando, Orlando, FL, USA.
Abstract:
PurposeLarge language models (LLMs) are increasingly used in clinical data workflows, but most clinicians lack the time, training, or support to build predictive models. We evaluated whether a ChatGPT-assisted workflow could enable clinicians to independently construct and execute a basic predictive model using routine perioperative data, using postoperative length of stay (LOS) in pediatric appendicitis as a feasibility use case.MethodsWe conducted a retrospective study of 228 children undergoing laparoscopic appendectomy. Ten routinely available preoperative variables were used. ChatGPT guided preprocessing and generated code for linear regression and random forest models using naïve and structured prompting. Models were trained using an 80/20 split and evaluated using MAE, RMSE, r, and R2.ResultsNaïve prompting performed poorly, while structured prompting improved model performance (MAE 0.34, RMSE 0.85, r 0.85). On the test set, linear regression achieved MAE 1.00 and random forest 0.77. Identified variables reflected expected clinical patterns. These findings demonstrate internal consistency of the workflow rather than new predictive insight.ConclusionsA structured ChatGPT workflow enabled clinicians to construct standard predictive models using routine data. These findings support feasibility of clinician-directed model development, but not clinical utility, of LLM-enabled workflows.