Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Regression Toward the Mean01:52

Regression Toward the Mean

6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Multiple Regression01:25

Multiple Regression

3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.9K
Weighted Mean00:57

Weighted Mean

6.2K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.2K
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

217
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
217
Choosing Between z and t Distribution01:25

Choosing Between z and t Distribution

3.5K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Working conditions and labour rights coverage of women domestic workers in Peru: a respondent-driven sampling study.

BMC public health·2026
Same author

The Relationships between physical activity, sedentary behaviour, sleep, and dementia: A systematic review and meta-analysis of cohort studies.

PloS one·2026
Same author

Investigating Early Kinetics in Plasma ctDNA and Peripheral T-cell Receptor Repertoire to Predict Treatment Outcomes to PD-1 Inhibitors in Head and Neck Squamous Cell Carcinoma.

Clinical cancer research : an official journal of the American Association for Cancer Research·2025
Same author

Community Engagement in Long Covid Research: Process, Evaluation and Recommendations From the Long COVID and Episodic Disability Study.

Health expectations : an international journal of public participation in health care and health policy·2025
Same author

Circumstances surrounding opioid toxicity deaths within shelters in Ontario, Canada, before and during the COVID-19 pandemic: a population-based descriptive cross-sectional study.

BMJ public health·2025
Same author

Lay rescuer intervention in fatal drownings in Canada, 2010-2019: a population-based cross-sectional analysis.

Injury prevention : journal of the International Society for Child and Adolescent Injury Prevention·2025

Related Experiment Video

Updated: Jan 4, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K

Unweighted regression models perform better than weighted regression techniques for respondent-driven sampling data:

Lisa Avery1,2, Nooshin Rotondi3,4, Constance McKnight5

  • 1York University, 4700 Keele St, Toronto, ON, M3J 1P3, Canada. lavery@maths.otago.ac.nz.

BMC Medical Research Methodology
|October 31, 2019
PubMed
Summary

For respondent-driven sampling (RDS) data, unweighted Poisson regression models are recommended for accurate risk estimation. Weighted regression models showed substantial bias and high error rates, making them less reliable for analyzing RDS data.

More Related Videos

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.6K
Watershed Planning within a Quantitative Scenario Analysis Framework
12:44

Watershed Planning within a Quantitative Scenario Analysis Framework

Published on: July 24, 2016

8.4K

Related Experiment Videos

Last Updated: Jan 4, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K
Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.6K
Watershed Planning within a Quantitative Scenario Analysis Framework
12:44

Watershed Planning within a Quantitative Scenario Analysis Framework

Published on: July 24, 2016

8.4K

Area of Science:

  • Epidemiology
  • Statistical Modeling

Background:

  • Respondent-driven sampling (RDS) is a method for collecting data from hidden populations.
  • The optimal statistical approach for analyzing RDS data, particularly regarding weighted versus unweighted regression, remains unclear.

Purpose of the Study:

  • To evaluate the validity of different regression models for analyzing data from respondent-driven sampling.
  • To assess the impact of weights and clustering controls on estimating the risk of group membership.

Main Methods:

  • Simulated 12 networked populations with varying homophily and prevalence using 1000 RDS samples each.
  • Applied weighted and unweighted binomial and Poisson general linear models with various clustering controls.
  • Evaluated models based on validity, bias, and coverage rates for prevalence estimation.

Main Results:

  • Unweighted log-link (Poisson) models maintained nominal type-I error rates.
  • Weighted binomial regression models exhibited substantial bias and unacceptably high type-I error rates.
  • RDS-weighted logistic regression provided the highest coverage rates for prevalence, except at low prevalence (10%) where unweighted models were preferred.

Conclusions:

  • Regression analysis of RDS data requires caution due to potential influence of reported degree.
  • Unweighted Poisson regression is recommended for analyzing RDS data to ensure reliable risk estimates.
  • Unweighted models are advised for prevalence estimation at low population prevalence levels.