Related Experiment Video
Updated: Jul 26, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Context Variance Evaluation of Pretrained Language Models for Prompt-based Biomedical Knowledge Probing
Zonghai Yao1, Yi Cao1, Zhichao Yang1
1College of Information and Computer Science, University of Massachusetts Amherst, Amherst, MA, USA.
This study enhances biomedical knowledge probing for pretrained language models (PLMs). New methods improve evaluation of complex relations, making BioLAMA more reliable for assessing PLM understanding.
Area of Science:
- Computational linguistics
- Biomedical informatics
- Artificial intelligence
Background:
- Pretrained language models (PLMs) are increasingly researched for their learned knowledge.
- Fill-in-the-blanks tasks, like cloze tests, are common for knowledge assessment.
- Existing benchmarks like BioLAMA face limitations due to prompt biases and data distribution.
Purpose of the Study:
- To address the unreliability and instability of prompt-based knowledge probing in BioLAMA.
- To develop improved methods for evaluating the factual knowledge of PLMs in the biomedical domain.
- To introduce a novel evaluation metric that better captures PLM understanding beyond simple recall.
Main Methods:
- Introduced context variance in prompt generation to mitigate probing biases.
- Proposed a new rank-change-based evaluation metric to assess PLM knowledge.
- Introduced the concept of 'Misunderstand' to differentiate true understanding from memorization.
Main Results:
- Experiments on 12 PLMs demonstrated that context variance prompts enhance BioLAMA's suitability for large-N-M and rare relations.
- The proposed Understand-Confuse-Misunderstand (UCM) metric proved effective in evaluating PLMs.
- Control experiments helped disentangle genuine understanding from simple data copying.
Conclusions:
- The developed context variance prompts and UCM metric offer a more robust and reliable evaluation of biomedical knowledge in PLMs.
- These advancements address key limitations of previous methods, particularly for complex and long-tailed biomedical data.
- The findings contribute to a deeper understanding of what knowledge PLMs truly acquire and how to effectively measure it.
Related Concept Videos
Improving Translational Accuracy
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...

