FSCME:一种特征选择方法,将相对应和最大信息系数与重相结合.
IEEE journal of biomedical and health informatics
|June 4, 2024
概括
特性选择方法由FSCME改进,它使用Copula相关性和最大信息系数来减少冗余并提高分类准确性. 这种方法为数据挖掘任务提供了更有效的功能子集.
科学领域:
- 数据挖掘 数据挖掘
- 机器学习 机器学习
- 信息理论 信息理论
背景情况:
- 在数据挖掘中,特征选择至关重要,但基于的方法可能是复杂和冗余的.
- 现有的方法难以准确地测量特征相关性和冗余性.
研究的目的:
- 引入FSCME,一种新的特征选择方法.
- 通过减少冗余和提高精度来解决基于的特征选择的局限性.
主要方法:
- 在FSCME中,Ccor用于冗余测量,最大信息系数 (MIC)用于相关性估计.
- 重法 (EWM) 为Ccor和MIC分配权重,以实现平衡的方法.
- 该方法考虑了功能标签的相关性和功能冗余性.
主要成果:
- 与其他六种方法相比,FSCME确定了一个更有效的特征子集.
- 拟议的方法显著提高了后续集群分类中的分类性能.
- 实验结果验证了FSCME在特征选择中的有效性.
结论:
- 在数据挖掘中,FSCME为特征选择提供了强大而有效的解决方案.
- 整合Ccor,MIC和EWM可以提高特征选择的可靠性和性能.
- 这种方法为改善数据分析和机器学习模型性能提供了有价值的工具.
更多相关视频
09:23Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
8.1K
07:12Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time
Published on: July 1, 2014
12.3K
相关概念视频
Coefficient of Correlation
6.1K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.1K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Extraction: Partition and Distribution Coefficients
2.4K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.4K
Calculating and Interpreting the Linear Correlation Coefficient
5.9K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
5.9K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
