Evaluation of Respondent-Driven Sampling Prevalence Estimators Using Real-World Reported Network Degree.
Lisa Avery1,2, Michael Rotondi3
1Department of Biostatistics, Princess Margaret Cancer Centre, University Health Network, Toronto, ON, Canada.
Summary
Respondent-driven sampling (RDS) provides prevalence estimates for hard-to-reach groups. Performance varies with network characteristics, requiring careful estimator selection for accurate disease and trait prevalence measurement.
Area of Science:
- Epidemiology
- Network Science
- Statistical Modeling
Background:
- Respondent-driven sampling (RDS) is a key method for estimating prevalence of traits or diseases in hidden or marginalized populations.
- Accurate prevalence estimation is crucial for public health interventions and resource allocation.
Purpose of the Study:
- To evaluate the performance of Respondent-driven sampling (RDS) estimators under diverse conditions.
- To assess the impact of network structure, trait prevalence, and homophily on RDS accuracy.
- To compare different RDS estimators in simulated and real-world network settings.
Main Methods:
- Utilized large simulated social networks (N=20,000) based on real-world RDS degree data.
- Employed an empirical Facebook network (N=22,470) for evaluating estimators.
- Assessed estimators for both binary and categorical trait prevalence.
Main Results:
- Prevalence estimate variability is higher with real-world network degrees compared to assumed Poisson distributions, leading to reduced coverage.
- Newer RDS estimators show good performance when sample size is a significant fraction of the population.
- Bias in prevalence estimates emerges when the total population size remains unknown.
Conclusions:
- The choice of the optimal Respondent-driven sampling (RDS) estimator is context-dependent.
- Study-specific considerations, including statistical properties and population knowledge, are vital for selecting the best RDS estimator.
- Understanding network characteristics is crucial for improving the reliability of RDS in prevalence studies.
Related Concept Videos
Testing a Claim about Population Proportion
3.4K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.4K
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Estimating Population Mean with Unknown Standard Deviation
8.2K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.2K
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Estimating Population Mean with Known Standard Deviation
8.8K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
8.8K


