Related Experiment Videos
Patient-Facing AI System for Symptom Guidance Using Simulated Encounters and Physician Review: Analytical Validation
Anil Patwardhan1, Nishant Verma1, Poulami Barman1
1Verily Life Sciences LLC (United States), Dallas, TX, United States.
Background:
Rapid advances in large language models (LLMs) have expanded interest in patient-facing health applications that support symptom assessment and care-seeking decisions. Although AI-enabled symptom-guidance tools could improve patient navigation and recognition of clinically serious conditions, inappropriate recommendations may result in missed needed care or unnecessary escalation. Structured predeployment evaluation is therefore needed before prospective clinical use.
Objective:
This formative predeployment study aimed to analytically validate a locked configuration of the Personal Health Assistant (PHA; Verily Health), an LLM-enabled patient-facing symptom-guidance tool, under controlled conditions using simulated patient encounters and expert physician review.
Methods:
We conducted a prospective predeployment analytical validation study using synthetic patient data. The analysis set included 772 synthetic cases generated using published telephone-triage protocols and persona-based, LLM-assisted case generation. Each case included a clinical vignette, structured medical history information, and a simulated multiturn patient conversation. All cases were reviewed by a panel of 5 independent physicians without access to PHA outputs in order to provide level-of-care recommendations. Case-level ground truth was derived from aggregated physician ratings using prespecified plurality and tie-handling rules. Urgent undertriage, nonurgent undertriage, and overtriage were evaluated against prespecified performance criteria. Secondary analyses assessed the clinical appropriateness of PHA-generated possible conditions and escalation, self-care, and laboratory-testing recommendations.
Results:
The final analysis set contained 772 cases, including 2 with unevaluable ground truth that were excluded from primary Endpoint denominators. Urgent undertriage was identified in 28 of 406 cases (6.9%, 95% CI 4.6%-9.8%), nonurgent undertriage in 56 of 305 (18.4%, 95% CI 14.2%-23.2%), and overtriage in 40 of 364 (11.0%, 95% CI 8.0%-14.7%). Overtriage met the prespecified criterion of less than 30%, whereas urgent and nonurgent undertriage did not meet their respective criteria of less than 5% and less than 15%, respectively. Secondary findings were generally supportive of the clinical appropriateness of the additional PHA outputs.
Conclusions:
PHA met the prespecified overtriage criterion but did not meet the urgent or nonurgent undertriage criteria. The findings identified undertriage as the principal performance gap requiring further product refinement. Because the study used synthetic cases, simulated conversations, and enriched urgency categories under controlled conditions, the results may not reflect performance in real-world populations or deployment settings. Prospective evaluation with actual users in the intended-use population is needed before broader clinical use.