Related Experiment Video
Updated: Oct 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
MedNLI-Pain benchmarks a minimally necessary requirement for large language model alignment
Christopher Wong1, Divy Kumar2, Amiin Muse3
1Department of Medicine, Rutgers Robert Wood Johnson Medical School, New Brunswick, NJ 08901, United States.
Objectives:
To simultaneously evaluate and compare LLM performance in clinical reasoning with a minimally necessary criteria for clinical alignment with patients.
Material And Methods:
We adapt the MedNLI dataset to create MedNLI-Pain, a new benchmark task that evaluates if LLMs that make accurate clinical inferences about patients can also anticipate when those patients are in pain. MedNLI-Pain consists of patient vignettes originating from MedNLI that are then annotated with physician assessments of the described patient's pain.
Results:
All LLMs had higher agreement with physicians in MedNLI than in MedNLI-Pain and its counterfactually augmented subset. This trend persisted when accounting for the higher difficulty of MedNLI-Pain by normalizing to human performance. Large language model performance improved with more training data in MedNLI but not MedNLI-Pain.
Discussion:
Large language models may lack necessary capabilities (eg, recognizing patients' pain) for clinical alignment. Success at a clinical reasoning task, in this case medical natural language inference, may not entail commensurate ability to align with patients' interests.
Conclusions:
Further research on alignment with patients is needed because the implications may direct future research into either closed-ended or agentic applications.