Related Experiment Video
Updated: May 23, 2025

09:05
Measurements of CO2 Fluxes at Non-Ideal Eddy Covariance Sites
Published on: June 24, 2019
7.8K
How different is different? Systematically identifying distribution shifts and their impacts in NER datasets.
Xue Li1, Paul Groth1
1Informatics Institute, University of Amsterdam, Science Park, Amsterdam, 1098 XH Netherlands.
Summary
Distribution shift significantly degrades natural language processing (NLP) model performance. This study systematically measures input and label shifts across NLP datasets, revealing performance drops and guiding model fine-tuning strategies.
Area of Science:
- Natural Language Processing (NLP)
- Machine Learning
Background:
- Distribution shift is a known challenge in NLP, where models trained on one data domain perform poorly on another.
- Existing research lacks systematic studies on detecting and quantifying the impact of distribution shifts on NLP task performance.
Purpose of the Study:
- To systematically detect and measure two types of distribution shift: input shift and label shift.
- To investigate the impact of these shifts on model performance across various NLP tasks and datasets.
- To provide insights into NLP model pipeline definition and fine-tuning requirements.
Main Methods:
- Detection and measurement of input and label shifts.
- Evaluation across three distinct data representations.
- Analysis on 12 benchmark Named Entity Recognition (NER) datasets.
Main Results:
- Both input and label shifts cause significant performance degradation in NLP models.
- A specific example shows a 63-point F1 score drop when fine-tuning on a broad dataset and testing on a narrower one.
- Shift measurement correlates with the amount of data needed for effective fine-tuning.
Conclusions:
- Quantifying distribution shift is crucial for understanding NLP model reliability.
- Shift measurement can inform decisions on whether a model is usable off-the-shelf or requires fine-tuning.
- Measuring distribution shift is vital for defining robust NLP model pipelines.
More Related Videos
Related Concept Videos
Distribution and Dispersion
21.5K
To understand intra-specific interactions in populations, scientists measure the spatial arrangement of species individuals. This geographic arrangement is known as the species distribution or dispersion. Highly territorial species exhibit a uniform distribution pattern, in which individuals are spaced at relatively equal distances from one another. Species that are highly tied to particular resources, such as food or shelter, tend to concentrate around those resources, and thus exhibit a...
21.5K
Steps in Outbreak Investigation
102
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
102
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K
Distribution Reliability and Automation
103
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
103
Types of Skewness
11.3K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
11.3K
Midrange
3.6K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
3.6K

