基于参数的随机森林获得了用于分类和回归问题的信息.
Vera Ignatenko1, Anton Surkov1, Sergei Koltcov1
1Social and Cognitive Informatics Laboratory, National Research University Higher School of Economics, Saint-Petersburg, Russia.
这项研究通过结合Renyi,Tsallis和Sharma-Mittal等变形度来增强随机森林算法,显著提高了对基准数据集的分类和回归任务的预测准确性.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 复杂的系统复杂的系统.
背景情况:
- 随机森林是很受欢迎的分类和回归,卓越表格数据.
- 传统的随机森林使用香农,可能会限制预测准确度.
- 来自复杂系统的扭曲为增强这些算法提供了一种新的方法.
研究的目的:
- 探索变形的有效性,以提高随机森林预测的准确性.
- 为随机森林引入基于Renyi,Tsallis和Sharma-Mittal的新型信息获取.
- 在各种分类和回归基准数据集上评估这些修改.
主要方法:
- 开发了使用Renyi,Tsallis和Sharma-Mittal的信息获取.
- 将这些新的信息获取集成到随机森林算法中.
- 在六个基准数据集 (三个分类,三个回归) 上测试了修改过的算法.
主要成果:
- 分类:Renyi变使准确度提高了19-96%,Tsallis提高了20-98%,Sharma-Mittal提高了22-111%.
- 回归:根据数据集,变形的值可以提高R2预测的2-23% .
- 所有测试的变形都超过了基于香农的经典随机森林.
结论:
- 扭曲的入量为随机森林准确性提供了相对于香农入量的显著改进.
- 这些新的方法为分类和回归任务提供了增强的预测能力.
- 这些发现表明,在使用复杂系统的概念来推进机器学习算法方面,这是一个有前途的方向.
更多相关视频
09:23Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
相关概念视频
Survival Tree
Building a Survival Tree
Constructing a...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Introduction to Nonparametric Statistics
One of...
Random Variables
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
