相关实验视频
Updated: Sep 12, 2025

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.2K
SRBench++:符号回归的原则性基准测试与领域专家解释
F O de Franca1, M Virgolin2, M Kommenda3
1Center for Mathematics, Computation and Cognition (CMCC), Heuristics, Analysis and Learning Laboratory (HAL), Federal University of ABC, Santo Andre, Brazil.
概括
这项研究通过包括专家解释性评估和分析各种数据科学任务的性能来增强符号回归 (SR) 基准测试. 它解决了当前基准标准的局限性,以实现更强大的算法评估.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 科学计算科学计算
背景情况:
- 符号回归 (SR) 旨在发现可解释的分析模型.
- 像SRBench这样的当前基准具有有限的解释性评估和特定任务的绩效评估.
- 模型大小对于真正的可解释性来说是一个不够的指标.
研究的目的:
- 为符号回归算法提出和评估一种改进的基准分析方法.
- 解决SRBench在评估可解释性和特定回归子任务方面的局限性.
- 为SR算法能力提供更全面的评估.
主要方法:
- 开发了一种新的基准分析方法,包括专家的解释性评估.
- 对不同数据科学任务属性的算法进行评估,包括特征选择和局部最小值规避.
- 评估了12个使用增强基准的现代符号回归算法.
主要成果:
- 在12个符号回归算法中确定了关键挑战和性能差异.
- 证明了专家驱动的解释性评估的价值.
- 强调需要制定基准,以捕捉特定数据科学挑战的绩效.
结论:
- 目前的符号回归基准需要加强全面的算法评估.
- 专家判断对于评估超出简单指标的模型解释性至关重要.
- 未来的基准标准应该包含多样化的任务属性,以更好地反映现实世界的适用性.
相关概念视频
Range Rule of Thumb to Interpret Standard Deviation
9.3K
The range rule of thumb in statistics helps us calculate a dataset's minimum and maximum values with known standard deviation. This rule is based on the concept that 95% of all values in a dataset lie within two standard deviations from the mean.
For instance, the range rule of thumb can be used to find the tallest and the shortest student in a class, given the mean student height and standard deviation. If the mean student height is 1.6 m and the standard deviation, s is 0.05 m, the height...
For instance, the range rule of thumb can be used to find the tallest and the shortest student in a class, given the mean student height and standard deviation. If the mean student height is 1.6 m and the standard deviation, s is 0.05 m, the height...
9.3K
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.3K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.3K
Empirical Method to Interpret Standard Deviation
5.4K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
5.4K
Testing a Claim about Standard Deviation
2.5K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.5K

