在多站点fMRI研究中克服站点变性:用于增强机器学习模型的自动编码框架
Fahad Almuqhim1,2, Fahad Saeed3
1Knight Foundation School of Computing and Information Sciences (KFSCIS), Florida International University, Miami, FL, USA.
Neuroinformatics
|September 2, 2025
概括
自动编码器有效地协调多站点功能磁共振成像 (fMRI) 数据,提高机器学习模型的概括性. 这种新的方法减少了特定地点的变化和数据泄露,提高了临床应用的可靠性.
科学领域:
- 神经成像
- 机器学习
- 数据协调
背景情况:
- 多站点功能磁共振成像 (fMRI) 数据协调对于可通用机器学习 (ML) 模型至关重要.
- 像ComBat这样的传统方法可能无法捕获复杂的非线性位置变化,并可能导致数据泄露.
- 这可能会限制在统一数据上训练的ML模型的可靠性和临床适用性.
研究的目的:
- 提出和评估自动编码器 (AE) 作为协调多站点fMRI数据的新方法.
- 利用AEs的非线性表示学习来减少特定地点的影响,同时保持生物特征.
- 解决传统统计协调技术中存在的数据泄露问题.
主要方法:
- 开发并实施使用各种自动编码器 (AE) 架构 (AE,SAE,TAE,DAE) 进行fMRI数据协调的框架.
- 评估了自闭症脑成像数据交换I (ABIDE-I) 数据集的框架 (1,035名受试者,17个中心).
- 根据基线方法对员工的绩效进行交叉验证.
主要成果:
- 所有AE变体都显示出与基线相比具有统计意义的改善 (p < 0. 01).
- 在LOSO交叉验证中,平均准确度的提高范围为3. 41%至5. 04%.
- 在保持神经生物学特征的同时,AE有效降低了特定位置的变异性.
结论:
- 自动编码器提供了一种强大的非线性方法来协调多站点fMRI数据.
- 这种方法提高了下游神经成像分析的稳定性和可重复性.
- 拟议的AE框架有效地减少了数据泄露,提高了临床使用的ML模型可靠性.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Variability: Analysis
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
Multiple Regression
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...


