更多可解释的特征选择表示在概率矩阵因子化中的零配编码.
Joshua C Chang1, Patrick A Fletcher1, Jungmin Han1
1National Institutes of Health & medεrrata.
概括
层次Poisson矩阵分解 (HPF) 缺乏编码器稀疏性,阻碍了可解释性. 本研究介绍了一种编码器分散的HPF,使用通用添加模型 (GAM) 来改善计数数据分析中的特征选择和可解释性.
科学领域:
- 医疗信息学 医疗信息学
- 计算生物学 计算生物学
- 统计建模 统计建模
背景情况:
- 减小尺寸对于可解释计数数据分析至关重要.
- 稀疏的概率非负矩阵分解 (NMF) 方法,如等级的波桑矩阵分解 (HPF),通过稀疏解码提供可解释性.
- 然而,HPF缺乏编码器稀疏性,限制了其定义因子-特征关系的能力.
研究的目的:
- 为了解决HPF中编码器稀疏性的缺陷.
- 开发一种强制执行编码器稀疏性的方法,以提高可解释性和功能选择.
- 为了证明编码器稀疏性在医疗信息学应用中的实际实用性.
主要方法:
- 在HPF框架内自始终强制执行编码器稀疏性.
- 使用通用添加模型 (GAM) 将表示坐标与原始数据特征联系起来.
- 将增强方法应用于模拟数据和Medicare患者共同疾病的真实数据集.
主要成果:
- 提出的方法成功地强制执行编码器稀疏性,与标准HPF不同.
- 该方法可以识别每个表示坐标的相关特征,方便特征选择.
- 在医疗保险患者的住院伴随性疾病中展示了实际应用.
结论:
- 在HPF中强制执行编码器稀疏性显著提高了模型解释性和特征选择能力.
- 将GAM集成为实现编码器稀疏性提供了一个强大的框架.
- 这种方法为分析医疗信息学及其他领域的高维数计数数据提供了有价值的工具.
更多相关视频
相关概念视频
Extraction: Partition and Distribution Coefficients
5.2K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
5.2K
Expected Frequencies in Goodness-of-Fit Tests
8.8K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
8.8K
Vector Algebra: Method of Components
20.2K
It is cumbersome to find the magnitudes of vectors using the parallelogram rule or using the graphical method to perform mathematical operations like addition, subtraction, and multiplication. There are two ways to circumvent this algebraic complexity. One way is to draw the vectors to scale, as in navigation, and read approximate vector lengths and angles (directions) from the graphs. The other way is to use the method of components.
In many applications, the magnitudes and directions of...
In many applications, the magnitudes and directions of...
20.2K
Gaussian Elimination: Problem Solving
230
Systems of linear equations in several variables are pivotal in modeling complex scenarios involving multiple unknowns and constraints. Such systems are widely used in various fields to represent relationships where several conditions must be simultaneously satisfied. Each variable in the system corresponds to an unknown quantity, while each equation imposes a linear constraint, leading to a structured approach for analyzing and solving real-world problems.A system of three equations with three...
230
Frequency-dependent Selection
24.3K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
24.3K
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K


