Related Experiment Video
Updated: Jun 18, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
A New Paradigm for High-dimensional Data: Distance-Based Semiparametric Feature Aggregation Framework via
Jinyuan Liu1, Xinlian Zhang2, Tuo Lin2
1Department of Biostatistics, Vanderbilt University, Nashville, Tennessee, U.S.A.
This study introduces a novel distance-based framework for analyzing high-dimensional data, preserving all features without selection. The approach uses semiparametric regression and U-statistics-based estimating equations for robust and efficient analysis.
Area of Science:
- Statistics
- Biostatistics
- Machine Learning
Background:
- High-dimensional data analysis presents challenges due to the curse of dimensionality.
- Traditional methods often rely on feature selection, leading to potential information loss.
- Existing inference methods may struggle with complex correlations in large datasets.
Purpose of the Study:
- To develop a distance-based framework for high-dimensional data analysis that avoids feature selection.
- To propose a semiparametric regression approach that encapsulates multiple high-dimensional variables.
- To introduce a robust and computationally feasible method for statistical inference.
Main Methods:
- A novel distance-based framework is proposed, focusing on pairwise outcomes of between-subject attributes.
- Semiparametric regression models are developed to handle multiple sources of high-dimensional variables.
- U-statistics-based estimating equations (UGEE) are utilized to address interlocking correlations, leveraging their unique efficient influence function (EIF).
Main Results:
- The proposed semiparametric estimators are robust to distributional misspecification.
- Achieved root-n consistency and asymptotic optimality facilitate reliable statistical inference.
- The framework effectively circumvents information loss associated with feature selection.
Conclusions:
- The developed approach enhances model interpretability and computational feasibility for high-dimensional data.
- It offers a powerful alternative to traditional methods, particularly for complex datasets like human microbiome and wearables data.
- The method preserves information and provides robust, efficient inference.
Related Concept Videos
Maximum Size of Aggregate
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Distance Measurements by Taping
Selected Data About Geographic Locations
Unsoundness of Aggregate due to Volume Change

