交叉验证选择对pBCI分类指标的影响:透明报告的教训
Felix Schroeder1, Stephen Fairclough1, Frederic Dehais2
1School of Psychology, Liverpool John Moores University, Liverpool, United Kingdom.
Frontiers in neuroergonomics
|July 16, 2025
概括
交叉验证选择在神经适应技术研究中显著偏差结果. 使用脑电图 (EEG) 数据的机器学习分类器的不同评估方法可以改变精度高达30.4%,影响可重现性.
科学领域:
- 神经科学是一个神经科学.
- 机器学习 机器学习
- 人与计算机的交互
背景情况:
- 神经适应技术是一种被动的大脑-计算机接口 (pBCI),利用神经生理信号进行人机交互.
- 评估机器学习和信号处理对于pBCI开发至关重要.
- 离线评估方法可以引入偏见,影响报告的准确性和结论.
研究的目的:
- 调查不同交叉验证方案如何在神经适应技术研究中偏向性能指标.
- 突出数据分割程序对分类器评估的影响.
- 强调需要对交叉验证方法进行详细报告.
主要方法:
- 分析了来自74名参与者的三个独立的脑电图 (EEG) n-back数据集.
- 交叉验证方案的比较,区分那些尊重和不尊重数据块结构.
- 计算性能差异的95%置信区间.
主要成果:
- 绩效指标和结论因交叉验证选择而有所不同.
- 根据评估方法,分类器的性能会有很大的差异.
- 里曼最小距离 (RMDM) 分类器的准确性可能会有高达12.7%的差异.
- 过器银行通用空间模式 (FBCSP) +线性差异分析 (LDA) 的准确性可以有高达30.4%的差异.
结论:
- 交叉验证的实施显著影响了pBCI研究中报告的准确性.
- 评估方法的差异可能会阻碍研究的可重复性.
- 对数据分割程序的详细报告对于透明和可重复的科学发现至关重要.
相关概念视频
Sensitivity, Specificity, and Predicted Value
673
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
673
Confidence Coefficient
7.9K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.9K
Receiver Operating Characteristic Plot
336
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
336
Decision Making: P-value Method
5.7K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.7K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Aggregates Classification
387
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
387


