Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

6.2K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
6.2K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

7.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
7.3K
Weighted Mean00:57

Weighted Mean

6.1K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

182
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
182
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

416
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
416
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

358
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
358

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Improving Variance and Confidence Interval Estimation in Small-Sample Propensity Score Analyses: Bootstrap Versus Asymptotic Methods.

Statistics in medicine·2026
Same author

Variance Estimation for Weighted Average Treatment Effects.

Statistics in biosciences·2026
Same author

Sample Size and Power Calculations With Win Measures Based on Hierarchical Endpoints.

Statistics in medicine·2025
Same author

Assessing racial disparities in healthcare expenditure using generalized propensity score weighting.

BMC medical research methodology·2025
Same author

Average treatment effect on the treated, under lack of positivity.

Statistical methods in medical research·2024
Same author

Evaluating analytic models for individually randomized group treatment trials with complex clustering in nested and crossed designs.

Statistics in medicine·2024

Related Experiment Video

Updated: Dec 14, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K

Propensity score weighting under limited overlap and model misspecification.

Yunji Zhou1,2, Roland A Matsouaka1,3, Laine Thomas1,3

  • 1Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA.

Statistical Methods in Medical Research
|July 23, 2020
PubMed
Summary

Inverse probability weighting (IPW) methods in non-randomized studies can be unstable due to positivity assumption violations. Alternative methods like overlap, matching, and entropy weights offer improved bias, variance, and coverage, especially with misspecified models.

Keywords:
Propensity scoreinverse probability weightslimited overlapmodel misspecificationoverlap weightstrimming

More Related Videos

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.8K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.5K

Related Experiment Videos

Last Updated: Dec 14, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.8K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.5K

Area of Science:

  • Epidemiology
  • Biostatistics
  • Causal Inference

Background:

  • Propensity score weighting methods, particularly inverse probability weighting (IPW), are crucial for adjusting confounding in non-randomized studies.
  • IPW relies on the positivity assumption, requiring propensity scores bounded away from 0 and 1, which is often violated in practice, leading to unstable estimates.

Purpose of the Study:

  • To compare the performance of IPW and IPW trimming against overlap weights, matching weights, and entropy weights.
  • To evaluate these methods under conditions of limited overlap and misspecified propensity score models.

Main Methods:

  • Extensive simulation studies were conducted.
  • The performance of different propensity score weighting methods was assessed based on bias, root mean squared error, and coverage probability.

Main Results:

  • Overlap weights, matching weights, and entropy weights consistently outperformed IPW.
  • These alternative methods demonstrated superior performance across various scenarios, particularly concerning bias, root mean squared error, and coverage probability.

Conclusions:

  • Alternative propensity score weighting methods (overlap, matching, entropy) are more robust than traditional IPW.
  • These methods provide more stable and reliable treatment effect estimates in the presence of positivity violations and model misspecification.