Related Experiment Video
Updated: Aug 10, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Robust Distance Correlation for Variable Screening
Tianzhou Ma1, Fan Yang2, Hongjie Ke1
1Department of Epidemiology and Biostatistics, University of Maryland, College Park, Maryland, USA.
This study introduces a robust distance correlation (RDC) screening method for heavy-tailed, high-dimensional data. The new approach effectively identifies key genes in complex datasets, outperforming existing methods.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- High-dimensional data analysis is crucial for scientific discovery.
- Traditional methods struggle with computational complexity and ultrahigh-dimensional data.
- Existing sure screening methods often fail to address heavy-tailed data characteristics.
Purpose of the Study:
- To develop a novel sure screening method for heavy-tailed, high-dimensional data.
- To enhance feature selection robustness in modern big data applications.
- To improve the identification of biologically relevant genes in cancer genomics.
Main Methods:
- Introduction of a new sure screening method utilizing robust distance correlation (RDC).
- Estimation of distance correlation robustly in the presence of heavy-tailed data.
- Development of a False Discovery Rate (FDR) control procedure using Reflection via Data Splitting (REDS).
Main Results:
- The proposed RDC-based screening method demonstrates superior performance over existing procedures in simulations with heavy-tailed data.
- The method successfully identified biologically meaningful genes predictive of MAPK1 protein expression in pancreatic cancer.
- Application to The Cancer Genome Atlas (TCGA) RNA-seq data confirmed the method's effectiveness.
Conclusions:
- The novel RDC screening method offers a robust solution for feature selection in high-dimensional, heavy-tailed datasets.
- This approach advances dimensionality reduction techniques for modern big data challenges.
- The method has significant implications for identifying key biomarkers in complex diseases like pancreatic cancer.
Related Concept Videos
Coefficient of Correlation
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the strength of the linear...
Correlation and Regression
Calculating and Interpreting the Linear Correlation Coefficient
Quantifying and Rejecting Outliers: The Grubbs Test
Correlation
Two variables, for example, a and b, are said to be positively correlated if both variables move in the same direction. In other words, a positive correlation exists between two variables, a and b, if:
Correlations
