Related Experiment Video
Updated: Sep 26, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
ProMem-agent: Procedural memory-augmented large language model agents for clinical trajectory reasoning
Qianying He1, Xuan Liu1, Jingquan Liu2
1Faculty of Data Science, City University of Macau, Avenida Padre Tomás Pereira, Taipa, Macao SAR, China.
Background And Objective:
Clinical large language model (LLM) agents can interpret current clinical context but have limited mechanisms for converting longitudinal experience into compact, reusable units. We developed ProMem-Agent, a procedural-memory framework that represents recurrent early intensive-care trajectories as provenance-linked observational patterns for retrospective mortality-risk estimation.
Methods:
Adult ICU stays with at least 24 hours of observable data were represented as six consecutive four-hour state-action-response intervals. Memory candidates were extracted exclusively from the MIMIC-IV ICU training cohort, linked to source events, consolidated by semantic clustering, and organized in a similarity graph. For each new patient, hybrid semantic and graph retrieval selected three memory cards. Comparators included conventional and longitudinal EHR models, direct and Chain-of-Thought LLM prompting, patient-level Case-RAG, token-matched Case-RAG, semantic-only Procedure-RAG, and first-24-hour SOFA as a clinically established severity reference. Evaluation included patient-level bootstrap testing, calibration analysis, external validation, retrieval and clustering sensitivity analyses, perturbation experiments, and blinded expert review.
Results:
The internal test cohort contained 6368 ICU stays with 11.9% mortality. ProMem-Agent achieved an F1-score of 0.608, AUC of 0.836, AUPRC of 0.481, and Brier score of 0.086. Relative to Procedure-RAG, the incremental differences were modest (AUC +0.013; AUPRC +0.029). Without memory reconstruction or external recalibration, AUC/AUPRC values were 0.823/0.402 on eICU, 0.831/0.429 on MIMIC-III, and 0.803/0.361 on HiRID. External calibration slopes were 0.88, 0.91, and 0.85, respectively, compared with 0.97 internally.
Conclusions:
Procedural abstraction accounted for a larger share of the observed gain than graph propagation, while external miscalibration and dataset shift limited transportability of absolute risk. ProMem-Agent should therefore be interpreted as a retrospective research framework for studying reusable clinical-trajectory representations, not as a clinically deployable decision-support system. Prospective, site-specific, clinician-in-the-loop evaluation remains necessary.
