Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

173
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
173
Relative Risk01:12

Relative Risk

333
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
333
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

7.3K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
7.3K
Estimating Population Standard Deviation01:26

Estimating Population Standard Deviation

3.1K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.1K
Confidence Interval for Estimating Population Mean01:25

Confidence Interval for Estimating Population Mean

8.0K
A point estimate of the population mean is obtained from a single sample. Such a point estimate does not represent a population well because it needs to account for variability in the population. Single point estimate can also be biased despite the sample being selected randomly. Thus, a point estimate is often unreliable. A confidence interval is needed to reduce this unreliability.
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
8.0K
Estimating Population Mean with Unknown Standard Deviation01:22

Estimating Population Mean with Unknown Standard Deviation

8.3K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
8.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

DR1-DR2 enrichment in HBeAg-negative chronic hepatitis B: integration hotspot or survivor signature?

Gut·2026
Same author

Impact of chronic obstructive pulmonary disease on the prognosis of patients with extensive-stage small-cell lung cancer treated with chemoimmunotherapy.

Translational lung cancer research·2026
Same author

Emergent Dynamical Kondo Coherence and Competing Magnetic Order in a Correlated Kagome Flat-Band Metal CsCr_{6}Sb_{6}.

Physical review letters·2026
Same author

Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement Learning.

Advances in neural information processing systems·2026
Same author

Phenylpropanoids from <i>Rhodiola fastigiata</i> (Hook.f. & Thomson) Fu (Crassulaceae) and their antiplasmodial activities.

Natural product research·2026
Same author

Prediction models for mortality in patients with acute on chronic liver failure: systematic review and critical appraisal.

Frontiers in medicine·2026

Related Experiment Video

Updated: Sep 7, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.2K

Measuring re-identification risk using a synthetic estimator to enable data sharing.

Yangdi Jiang1,2, Lucy Mosquera2, Bei Jiang1

  • 1Department of Mathematical and Statistical Sciences, University of Alberta, Edmonton, Canada.

Plos One
|June 17, 2022
PubMed
Summary

A new risk estimator accurately measures re-identification risk in de-identified health data. This method, averaging two copula estimators, enhances privacy protection for secondary data analysis.

More Related Videos

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
05:35

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome

Published on: September 20, 2022

3.8K

Related Experiment Videos

Last Updated: Sep 7, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.2K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
05:35

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome

Published on: September 20, 2022

3.8K

Area of Science:

  • Health Informatics
  • Data Privacy
  • Statistical Modeling

Background:

  • De-identifying health data is crucial for secondary analysis while adhering to privacy regulations.
  • Existing re-identification risk metrics lack accuracy for sample-to-population attacks.
  • An accurate risk estimator is needed for adversaries matching microdata samples to population data.

Purpose of the Study:

  • Develop and evaluate an accurate risk estimator for the sample-to-population attack scenario.
  • Improve methods for assessing re-identification risk in de-identified datasets.

Main Methods:

  • Developed a novel risk estimator using synthetic population datasets.
  • Evaluated estimator accuracy via simulations on four datasets, comparing Gaussian and d-vine copula models.
  • Assessed performance against three existing risk estimation methods.

Main Results:

  • The average of two copula estimators achieved a median error below 0.05 across various sampling fractions.
  • This combined estimator significantly outperformed existing methods in accuracy.
  • Sensitivity analysis provided guidance for practical application, including de-identifying a COVID-19 survey dataset.

Conclusions:

  • Averaging two copula-based estimators provides a highly accurate method for re-identification risk assessment.
  • This approach offers a reliable foundation for managing privacy risks in de-identified health data sharing.