相关实验视频
Updated: Jun 25, 2025

07:42
A Data-Driven Approach to Quantifying Immune States in Sepsis
Published on: February 7, 2025
167
用全球COVID-19趋势和影响调查 (CTIS) 的疫苗接种数据,探索各种估计的大数据悖论
Youqi Yang1, Walter Dempsey1, Peisong Han2
1Department of Biostatistics, University of Michigan, Ann Arbor, MI, USA.
Science advances
|May 31, 2024
概括
像COVID-19趋势和影响调查 (CTIS) 这样的大型非概率调查显示,与较小的概率调查相比,在估计COVID-19疫苗接种率时,错误率更高. 这突出了大数据研究中潜在的选择偏见问题.
科学领域:
- 流行病学 流行病学
- 调查方法 调查方法
- 生物统计学 生物统计学
背景情况:
- 在非概率样本中的选择偏差使准确的统计推理复杂化.
- 评估COVID-19疫苗接种率需要可靠的数据,这往往受到采样方法的挑战.
研究的目的:
- 使用国家基准数据,比较来自大型非概率调查 (CTIS) 和小概率调查 (CVoter) 的COVID-19疫苗接种率估计的准确性.
- 评估CTIS在估计疫苗接种时间和子组差异方面的准确性.
主要方法:
- 从CTIS和CVoter对第一剂COVID-19疫苗接种率的估计错误与COVID疫苗情报网络基准数据进行比较.
- 分析了CTIS的平均平方误差,用于接种疫苗的连续和子组差异.
主要成果:
- 对于整体疫苗接种率,CTIS显示的平均估计误差 (0.37) 比CVoter (0.14) 更大.
- 使用CTIS估计差异 (随着时间的推移或子组之间的差异) 与总比率相比,有效样本大小增加.
- 大数据悖论可能超越美国,并根据具体估计而有所不同.
结论:
- 像CTIS这样的非概率调查可能会出现显著的估计错误,这凸显了选择偏差的持续挑战.
- 虽然CTIS对整体率显示出更高的误差,但分析差异可能会提高其实用性,但建议谨慎.
- 调查结果表明,大数据悖论并不仅限于美国,其影响是估计和依赖的.
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
353
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
353
Steps in Outbreak Investigation
122
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
122
Bias in Epidemiological Studies
242
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
242
Types of Skewness
11.5K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
11.5K
Causality in Epidemiology
385
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
385
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K

