Related Experiment Video
Updated: May 23, 2025

Measurements of CO2 Fluxes at Non-Ideal Eddy Covariance Sites
Published on: June 24, 2019
How different is different? Systematically identifying distribution shifts and their impacts in NER datasets
Xue Li1, Paul Groth1
1Informatics Institute, University of Amsterdam, Science Park, Amsterdam, 1098 XH Netherlands.
Abstract:
When processing natural language, we are frequently confronted with the problem of distribution shift. For example, using a model trained on a news corpus to subsequently process legal text exhibits reduced performance. While this problem is well-known, to this point, there has not been a systematic study of detecting shifts and investigating the impact shifts have on model performance for NLP tasks. Therefore, in this paper, we detect and measure two types of distribution shift, across three different representations, for 12 benchmark Named Entity Recognition datasets. We show that both input shift and label shift can lead to dramatic performance degradation. For example, fine-tuning on a wide spectrum dataset (OntoNotes) and testing on an email dataset (CEREC) that shares labels leads to a 63-points drop in F1 performance. Overall, our results indicate that the measurement of distribution shift can provide guidance to the amount of data needed for fine-tuning and whether or not a model can be used "off-the-shelf" without subsequent fine-tuning. Finally, our results show that shift measurement can play an important role in NLP model pipeline definition.
More Related Videos
Related Concept Videos
Distribution and Dispersion
Steps in Outbreak Investigation
Distributions to Estimate Population Parameter
Distribution Reliability and Automation
Types of Skewness
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
Midrange
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...

