在线推论高维通用线性模型与数据流的在线推论
Lan Luo1, Ruijian Han2, Yuanyuan Lin3
1Department of Biostatistics and Epidemiology, Rutgers School of Public Health, New Jersey, USA.
概括
本研究介绍了一种在线统计推理方法,用于使用流数据的高维通用线性模型. 新的在线 debiased lasso有效地估计回归系数,并使实时统计推理成为可能.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 高维通用线性模型 (GLM) 对于分析复杂数据集至关重要.
- 流数据需要有效的在线方法来实时估计和推断.
- 传统的离线方法不适合连续的数据流.
研究的目的:
- 开发一个在线统计推断方法,用于高维的GLMs与流数据.
- 为了实现实时估计和推断不断到达的数据.
- 在动态数据环境中解决离线方法的局限性.
主要方法:
- 提议一种针对数据流量量身定制的在线 debiased lasso 方法.
- 更新信心区间,仅使用历史总结统计数据.
- 加入一个额外的术语来纠正在线更新期间累积的近似错误.
主要成果:
- 在GLM中提出的在线debiased估计器被证明是异常正常的.
- 这种非对称的正常性为实时中间推理提供了理论上的理由.
- 数字实验验证了在线 debiased lasso 方法的有效性和性能.
结论:
- 开发的在线 debiased lasso 方法为使用高维流数据进行统计推理提供了强大的解决方案.
- 理论结果支持该方法用于实时分析的应用.
- 这种方法被证明是有效的,即使是在大规模的文本数据集上.
相关概念视频
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
64
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
64
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
433
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
433
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
117
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
117
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Statistical Methods for Analyzing Epidemiological Data
336
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
336
Friedman Two-way Analysis of Variance by Ranks
170
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
170


