Related Experiment Video
Updated: Jul 10, 2025

08:25
Polar Histogram Visualization of Acute Stress Disorder Scale Scores for Comprehensive Clinical Assessment
Published on: December 6, 2024
391
Effect of data harmonization of multicentric dataset in ASD/TD classification
Giacomo Serra1,2, Francesca Mainas3,4, Bruno Golosio1,2
1Department of Physics, University of Cagliari, Cagliari, Italy.
Brain Informatics
|November 25, 2023
Summary
Harmonizing neuroimaging data using the entire dataset improves Autism Spectrum Disorder (ASD) classification but causes data leakage. Internal harmonization, using only training data, prevents leakage while maintaining performance.
Area of Science:
- Neuroimaging
- Machine Learning
- Developmental Disorders
Background:
- Machine learning (ML) is crucial for analyzing neuroimaging data, especially Magnetic Resonance Imaging (MRI), to identify brain patterns in neurological disorders.
- Multicenter neuroimaging datasets are essential for ML model training but introduce biases due to site-specific variations.
- ComBat harmonization is a common technique to correct batch effects, yet it risks data leakage when applied to the entire dataset.
Purpose of the Study:
- To evaluate the impact of different data harmonization strategies on the classification of Autism Spectrum Disorders (ASD) using structural and functional MRI data.
- To compare external harmonization (whole dataset), internal harmonization (training set only), and no harmonization in the context of ASD detection.
- To determine if improved performance from external harmonization is due to data leakage.
Main Methods:
- Utilized structural and functional MRI data from the Autism Brain Imaging Data Exchange (ABIDE) dataset.
- Compared three approaches: external harmonization (before train/test split), internal harmonization (on training set only), and no harmonization.
- Assessed classification performance for Autism Spectrum Disorders (ASD) versus Typical Developing (TD) controls.
Main Results:
- External harmonization (whole dataset) yielded higher classification performance for both structural and connectivity features.
- Non-harmonized data and internal harmonization (training set only) showed comparable, lower performance.
- The superior performance of external harmonization was attributed to data leakage, not sample size for model estimation.
Conclusions:
- External harmonization, while boosting performance, introduces data leakage, potentially inflating results.
- To prevent data leakage and ensure reliable model generalization, harmonization models should be trained exclusively on the training dataset.
- Internal harmonization is recommended for robust ML analysis of multicenter neuroimaging data in ASD research.
Related Concept Videos
How Data are Classified: Categorical Data
33.1K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
33.1K
Aggregates Classification
327
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
327
Classification of Systems-I
188
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
188
Classification of Systems-II
149
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
149
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K

