Related Experiment Video
Updated: Jun 20, 2026

Gene Regulation and Targeted Therapy in Gastric Cancer Peritoneal Metastasis: Radiological Findings from Dual Energy CT and PET/CT
Published on: January 22, 2018
Rules-Augmented GLM-5.1 Prompting for Four-Class Chest CT Protocol Selection
1Schulich School of Medicine and Dentistry, Western University, 1151 Richmond St, London, ON, N6A 5C1, Canada. kartikg9@gmail.com.
None:
Rule-augmented large language models (LLMs) may support radiology protocol selection, but their value in small, imbalanced local tasks remains uncertain. We evaluated whether GLM-5.1 prompting remained competitive with classical text classifiers for four-class chest CT protocol selection when tested on held-out cases. We retrospectively analyzed 755 chest CT requests from a single Canadian center (January 2022-January 2023). Reference labels were the clinically performed protocols: chest with contrast (n = 564), chest without contrast (n = 97), interstitial/high-resolution CT for interstitial lung disease (n = 27), and low-dose CT (n = 67). Models were compared on stratified 70:30 held-out splits: majority baseline, random forest with TF-IDF features and random oversampling (RF-ROS), fine-tuned BioClinicalBERT, and GLM-5.1 with a classification-rules prompt. Primary metrics were balanced accuracy and macro-F1; paired held-out predictions were compared with McNemar testing. On the primary split (n = 227; seed 42), the majority baseline achieved accuracy 0.749 and balanced accuracy 0.250. RF-ROS achieved accuracy 0.907, balanced accuracy 0.793, and macro-F1 0.815; BioClinicalBERT achieved 0.881, 0.759, and 0.753; and GLM-5.1 achieved 0.916, 0.918, and 0.838, respectively. GLM-5.1 was not significantly different from RF-ROS (p = 0.82) or BioClinicalBERT (p = 0.20). It classified all 8 held-out Interstitial cases correctly. Rule-augmented GLM-5.1 prompting was feasible and not significantly different from imbalance-aware classical comparators on paired testing in this narrow single-center task. Results support local rule encoding as an adaptation strategy, not unsupervised replacement of classical ML or human verification.
