多中心数据集数据协调在ASD/TD分类中的影响
Giacomo Serra1,2, Francesca Mainas3,4, Bruno Golosio1,2
1Department of Physics, University of Cagliari, Cagliari, Italy.
Brain informatics
|November 25, 2023
概括
使用整个数据集协调神经成像数据可以改善自闭症谱系障碍 (ASD) 的分类,但会导致数据泄露. 内部协调,仅使用训练数据,防止泄漏,同时保持性能.
科学领域:
- 神经成像是一种神经成像.
- 机器学习 机器学习
- 发育障碍 发育障碍 发展障碍
背景情况:
- 机器学习 (ML) 对于分析神经成像数据至关重要,特别是磁共振成像 (MRI),以识别神经疾病中的大脑模式.
- 多中心神经成像数据集对于ML模型培训至关重要,但由于特定站点的变化而引入偏差.
- 康巴统一是纠正批量效应的常见技术,但当应用于整个数据集时,可能会有数据泄露的风险.
研究的目的:
- 用结构和功能MRI数据评估不同数据协调策略对自闭症谱系障碍 (ASD) 的分类的影响.
- 为了比较外部协调 (整个数据集),内部协调 (仅培训集) 和在ASD检测的背景下没有协调.
- 确定外部协调所带来的性能改善是否是由于数据泄露.
主要方法:
- 利用自闭症脑成像数据交换 (ABIDE) 数据集中的结构和功能MRI数据.
- 我们比较了三个方法:外部协调 (在训练/测试分割之前),内部协调 (仅在训练集上) 和没有协调.
- 评估了自闭症谱系障碍 (ASD) 与典型发育 (TD) 控制对象的分类性能.
主要成果:
- 外部协调 (整个数据集) 给结构和连接特征带来了更高的分类性能.
- 非协调数据和内部协调 (仅培训组) 显示了相似的,较低的表现.
- 外部协调的优异性能归因于数据泄露,而不是模型估计的样本大小.
结论:
- 外部协调,虽然提高了性能,但引入了数据泄露,可能会膨胀结果.
- 为了防止数据泄露,并确保可靠的模型通用化,协调模型应仅在训练数据集上进行训练.
- 建议内部协调,以便在ASD研究中对多中心神经成像数据进行强大的ML分析.
相关概念视频
How Data are Classified: Categorical Data
33.1K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
33.1K
Aggregates Classification
327
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
327
Classification of Systems-I
188
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
188
Classification of Systems-II
149
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
149
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K


