Related Experiment Video
Updated: May 13, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Clinical-ShiftEval: a framework for simulating and evaluating model adaptation in dynamic clinical NLP tasks
Fabián Villena1,2, Felipe Bravo-Marquez3, Jocelyn Dunstan4,5
1School of Dentistry, Pontificia Universidad Católica de Chile, Santiago, Chile.
Clinical natural language processing (NLP) models degrade with evolving data. A new framework, Clinical-ShiftEval, shows hybrid and in-context learning methods significantly improve model robustness against these real-world clinical shifts.
Area of Science:
- Clinical Natural Language Processing (NLP)
- Machine Learning in Healthcare
- Electronic Health Records (EHR) Data Analysis
Background:
- Clinical NLP models are crucial for extracting information from EHRs, but often fail in dynamic clinical environments.
- Evolving data distributions (e.g., new guidelines, diseases) cause significant performance degradation in existing models.
- Current evaluation methods assume static data, failing to capture real-world model adaptation challenges.
Purpose of the Study:
- To introduce Clinical-ShiftEval, a framework for simulating and evaluating NLP model adaptation in dynamic clinical settings.
- To assess model performance under label-set incompatibility (LSI) and task definition evolution (TDE) using real clinical text.
- To compare the effectiveness of continual training, in-context learning, and hybrid approaches for model adaptation.
Main Methods:
- Developed Clinical-ShiftEval to operationalize LSI and TDE through controlled data transitions from real datasets.
- Applied the framework to the Chilean Waiting List Corpus, modeling LSI with referral specialty classification and TDE with disease prioritization.
- Systematically compared continual training, in-context learning with large language models, and a hybrid adaptation strategy.
Main Results:
- Clinical-ShiftEval reliably simulates realistic clinical changes and produces interpretable performance drops, validating its utility for benchmarking.
- Conventional supervised models degraded significantly (up to 82% F1 drop for LSI, 43% for TDE).
- In-context learning reduced performance drops (35% LSI, 10% TDE), while a hybrid method achieved ~8% drop; continual training required >30% new data to surpass these.
Conclusions:
- Clinical-ShiftEval provides a robust method for benchmarking NLP model adaptation in dynamic EHR environments.
- In-context learning and hybrid approaches offer substantial resilience to evolving clinical data without needing new labeled data.
- Continual training with moderate new data is most effective for optimal accuracy recovery, eventually outperforming other methods.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Impression Management Techniques IV: Altercasting
Pharmacodynamic Models: Overview
Statistical Software for Data Analysis and Clinical Trials
Clearance Models: Physiological Models
The organ's clearance rate depends on the blood flow to the organ and the extraction ratio (E). The extraction ratio describes the organ's proficiency in drug...