Related Experiment Videos
Safety-aware AI for NSCLC trial pre-screening: a comparative proof-of-concept study of rule-based, single-agent, and
1Faculty of Health Sciences, University of Macau, Faculty of Health Sciences Building, E12 Avenida da Universidade, Taipa, Macau, 999078, China.
Background:
Non-small cell lung cancer (NSCLC) trials often involve complex biomarker-driven and line-specific eligibility criteria, making pre-screening labor-intensive and error-prone. In this high-stakes setting, false inclusion-incorrectly judging an ineligible patient as eligible for downstream review-poses a greater clinical risk than other error types.
Objective:
To compare safety-aware AI configurations for NSCLC trial pre-screening, with explicit focus on reducing false inclusion rates and supporting uncertainty-aware abstention under evidence limitations.
Methods:
We performed a retrospective proof-of-concept evaluation using interventional NSCLC trial eligibility text and structured synthetic patient summaries covering variations in stage, biomarker status, prior therapy, ECOG performance status, and exclusion-relevant factors. A total of 120 patient-trial pairs were manually reviewed and split into a development set (n = 30) and held-out test set (n = 90). Five configurations were compared: rule-based baseline, rule-based with conservative Safety Agent gating, single-agent LLM (GPT-4o-mini), and two lightweight multi-agent variants (M1, M2). The primary safety metric was false inclusion rate; secondary metrics were accuracy and uncertain rate. Secondary evaluations included a TCGA-derived stress test under incomplete evidence and a real-text external validation on 72 patient-trial pairs from PMC Open Access NSCLC case reports.
Results:
Safety Agent gating reduced the false inclusion rate from 11.63% to 1.16% in the combined 120-pair cohort (McNemar's test, p = 0.002) and from 12.12% to 1.52% in the held-out set, accompanied by an increase in uncertain rate. The single-agent LLM (GPT-4o-mini) achieved the highest observed accuracy (0.74, 95% CI 0.66-0.81), with no observed false inclusion among 86 gold-standard ineligible pairs (95% CI 0.0-4.3%) and an uncertain rate of 7.5%. Multi-agent M1 and M2 also achieved zero observed false inclusion on the structured cohort with distinct uncertainty profiles (25.8% and 1.7% uncertain, respectively); in the PMC real-text validation, M1 maintained zero false inclusion (0/44; 95% CI 0.0-8.0%) while M2 produced one false inclusion arising from an anomalous Safety Agent label upgrade on ambiguous unstructured text. A multi-provider evaluation across five LLM configurations spanning four providers found four of five achieving zero observed false inclusion on the full cohort. In the TCGA-derived stress test with missing biomarker and ECOG data, all evaluated systems showed zero observed false inclusion.
Conclusions:
In this proof-of-concept evaluation under controlled synthetic-input conditions, conservative safety gating substantially reduced false inclusion in rule-based pre-screening, and all LLM-based configurations achieved zero or near-zero false inclusion on the structured synthetic cohort. These findings support explicit false inclusion monitoring, uncertainty-aware abstention, and multi-dimensional evaluation as useful design principles for AI-assisted oncology trial pre-screening. However, results are limited to structured synthetic data, a single LLM family, and small manually reviewed cohorts; prospective multi-site validation on diverse unstructured EHR documentation is needed before safety or utility claims can be generalized.