学习了评估条件和相互信息的实用指南,以发现响应与共变动态的主要因素
Ting-Li Chen1, Hsieh Fushing2, Elizabeth P Chou3
1Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan.
Entropy (Basel, Switzerland)
|July 8, 2023
概括
本研究引入了使用分类数据动态的新数据分析框架. 它使用信息理论来识别统计模型中的关键因素,为复杂的数据挑战提供实用指南.
科学领域:
- 统计 统计 统计 统计
- 信息理论 信息理论
- 数据分析 数据分析
背景情况:
- 传统的统计方法通常依赖于明确的功能结构.
- 分析复杂的参数统计学主题可能具有挑战性.
- 发现数据中的潜在因素对于有效分析至关重要.
研究的目的:
- 将统计学主题重新构成一个响应与共变量 (Re-Co) 动态框架.
- 通过仅使用分类数据发现主要因素来解决数据分析任务.
- 开发使用信息理论测量方法进行因素选择的计算指南.
主要方法:
- 使用香农条件 (CE) 和相互信息 (I[Re;Co]) 进行因子选择.
- 使用分类探索性数据分析 (CEDA) 范式.
- 在应急表平台上评估信息理论测量.
主要成果:
- 根据[C1:可确认]标准,制定了评估CE和I[Re;Co]的实用准则.
- 开发方法来减轻在应急表中的维度诅咒的影响.
- 成功地将框架应用于六个Re-Co动态的例子,并扩展了场景.
结论:
- 拟议的Re-Co动态框架为统计数据分析提供了一种新的方法.
- 信息理论测量为分类数据中主要因素选择提供了有效的工具.
- 制定的指导方针促进了实用和高效的数据分析,即使对于高维数据集.
更多相关视频
09:23Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
8.1K
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
3.4K
相关概念视频
Standard Entropy Change for a Reaction
20.6K
Entropy is a state function, so the standard entropy change for a chemical reaction (ΔS°rxn) can be calculated from the difference in standard entropy between the products and the reactants.
20.6K
Introduction To Survival Analysis
289
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
289
Correlation of Experimental Data
256
Dimensional analysis simplifies complex physical problems and guides experimental investigations, but it does not provide complete solutions. It identifies the dimensionless groups that influence a phenomenon, but experimental data is needed to establish the specific relationships and validate theoretical predictions.
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
256
Statistical Methods for Analyzing Epidemiological Data
425
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
425
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
