Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Absolute and Local Extreme Values01:22

Absolute and Local Extreme Values

The highest and lowest values of a function, relative to a reference axis, are known as extreme values. These include absolute maximum and absolute minimum values, which represent the highest and lowest points the function reaches across its entire domain. Within a restricted portion of the function, the highest and lowest values are referred to as local maximum and local minimum values, respectively.Periodic functions, such as sine and cosine, show extreme values at infinitely many points due...
Unusual Results01:16

Unusual Results

Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ  from the mean, μ  is considered unusual.
Maximum unusual value = μ + 2σ
Minimum unusual value...
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

LLMs for Healthcare Process Orchestration: Promises and Challenges.

Studies in health technology and informatics·2026
Same author

Challenges of Bringing Augmented Reality into Nursing Medication Preparation.

Studies in health technology and informatics·2026
Same author

Modeling Blood Pressure Monitoring in openEHR: An Implementation Case Study.

Studies in health technology and informatics·2026
Same author

Modeling Reimbursement-Relevant Drug Data as a Property Graph.

Studies in health technology and informatics·2026
Same author

LLMs in Problem-Based Learning: The Grounding Issue.

Studies in health technology and informatics·2026
Same author

Large language models as cognitive shortcuts: a systems-theoretic reframing beyond bullshit.

Frontiers in artificial intelligence·2026

Related Experiment Video

Updated: Jun 4, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

Controlling false match rates in record linkage using extreme value theory.

Murat Sariyar1, Andreas Borg, Klaus Pommerening

  • 1Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center of the Johannes Gutenberg University Mainz, Germany. murat.sariyar@unimedizin-mainz.de

Journal of Biomedical Informatics
|March 1, 2011
PubMed
Summary

This study introduces a novel method using Extreme Value Theory (EVT) to estimate false match rates in record linkage. This approach reduces costs by eliminating the need for training data, improving data quality in medical research.

Related Experiment Videos

Last Updated: Jun 4, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

Area of Science:

  • Data Science
  • Biostatistics
  • Information Science

Background:

  • Data cleansing is crucial for high-quality data in disease registries and medical research.
  • Record linkage methods minimize errors like synonyms and homonyms, enhancing data integrity.
  • Homonym errors (false matches) are critical, where distinct entities are incorrectly identified as identical.

Purpose of the Study:

  • To present a new approach for estimating false match rates in record linkage.
  • To leverage Extreme Value Theory (EVT) for a more efficient and cost-effective solution.
  • To address the limitations of manual clerical review and existing statistical models requiring training data.

Main Methods:

  • Utilizing Extreme Value Theory (EVT) within the Fellegi and Sunter framework.
  • Applying the generalized Pareto distribution and mean excess plots for analysis.
  • Developing a method that does not require training data for threshold determination.

Main Results:

  • The proposed EVT-based approach effectively estimates false match rates.
  • The method significantly reduces costs by eliminating the need for training data.
  • Experimental results show comparable accuracy to methods with match status information.

Conclusions:

  • Extreme Value Theory offers a viable and cost-effective alternative for estimating false match rates.
  • This method enhances data quality in critical applications like medical research networks.
  • The approach provides a significant advantage by removing the dependency on calibration data.