相关实验视频
Updated: May 23, 2025

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.0K
对共变量分布模型的信息理论评估
Niklas Hartung1, Aleksandra Khatova2,3
1Institute of Mathematics, University of Potsdam, Karl-Liebknecht-Str. 24-25, 14476, Potsdam, Germany. niklas.hartung@uni-potsdam.de.
Journal of pharmacokinetics and pharmacodynamics
|March 28, 2025
概括
非高斯统计模型,包括连锁方程 (MICE) 的配方和多重归算,在生命科学中的复杂共变量分布中优于高斯模型. 这些先进的方法改善了虚拟人口的产生和缺失数据的归算.
科学领域:
- 统计建模 统计建模
- 生命科学 生命科学
- 数据科学数据科学数据科学
背景情况:
- 在生命科学中的共变量分布往往表现出非高斯边缘和复杂的非线性相关性.
- 标准的多变量高斯模型无法准确地表示这些复杂的结构.
- 现有的非高斯框架,如形模型和链式方程 (MICE) 的多重归算,缺乏系统的适合性评估.
研究的目的:
- 系统地评估非高斯共变量分布模型的适合性.
- 为了比较基于copula的模型和MICE与更简单的近似模型的性能.
- 在不同的数据条件下评估模型性能,包括大小,维度和缺失.
主要方法:
- 利用Kullback-Leibler (KL) 差异作为一个规模不变的合适性标准.
- 开发了一种用于KL分歧置信区间的新方法,使用最近邻近估计器和亚抽样.
- 应用于各种数据集的方法,具有连续和离散的共变量,大小和维度各不相同.
主要成果:
- 非高斯模型始终表现出高斯或尺度变换近似相比较优异的KL分歧.
- 对于隐性变量和大量缺失数据,KL差异估计结果证明是可靠的.
- 哥普拉模型对新数据进行了更好的概括,而MICE显示出过度拟合的倾向.
- 参数模型和MICE在数据集维度上比非参数模型更好地进行缩放.
结论:
- 非高斯模型为建模现实的生命科学共变量分布提供了显著的改进.
- 由于它们的概括能力,建议使用Copula模型,而MICE则需要对测试数据进行仔细评估.
- 参数配方模型和MICE为高维数据集提供可扩展的解决方案.
相关概念视频
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K
Probability Distributions
6.7K
The probability of a random variable x is the likelihood of its occurrence. A probability distribution represents the probabilities of a random variable using a formula, graph, or table. There are two types of probability distribution– discrete probability distribution and continuous probability distribution.
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
6.7K
Friedman Two-way Analysis of Variance by Ranks
130
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
130
Binomial Probability Distribution
10.2K
A binomial distribution is a probability distribution for a procedure with a fixed number of trials, where each trial can have only two outcomes.
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
The outcomes of a binomial experiment fit a binomial probability distribution. A statistical experiment can be classified as a binomial experiment if the following conditions are met:
There are a fixed number of trials. Think of trials as repetitions of an experiment. The letter n denotes the number of trials.
There are only two possible outcomes,...
10.2K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
113
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
113

