Related Experiment Video
Updated: Jun 7, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Multiple feature selection based on an optimization strategy for causal analysis of health data
Ruichen Cong1, Ou Deng1, Shoji Nishimura2
1Graduate School of Human Sciences, Waseda University, 2-579-15 Mikajima, Tokorozawa, 359-1192 Saitama Japan.
This study introduces a novel framework for health data analysis, utilizing an optimization strategy for multiple feature selection to identify key health indicators. The approach effectively reduces data complexity and validates causal relationships, enhancing healthcare insights.
Area of Science:
- Health Informatics
- Data Science
- Biostatistics
Background:
- Advancements in information technology and wearable devices generate complex health datasets.
- Causal graphs are crucial for understanding health feature relationships but face computational challenges.
- Feature selection is vital for managing high-dimensional health data and improving analytical efficiency.
Purpose of the Study:
- To present a framework for multiple feature selection using an optimization strategy for causal analysis of health data.
- To address challenges in health analytics, including large feature sets and computational demands.
- To enhance the identification of significant relationships within complex health data.
Main Methods:
- A Weighted Total Score (WTS) index was defined to assess combined feature importance from multiple selection methods.
- A greedy algorithm was integrated to optimize weights for each feature selection method.
- Causal graphs were constructed using selected features, with statistical significance assessed for causal paths.
Main Results:
- The proposed framework reduced the number of features while improving model performance compared to baseline models.
- The statistical significance of feature relationships identified through causal graphs was validated.
- The approach demonstrated effectiveness on both a custom dataset and an open diabetes dataset.
Conclusions:
- The developed framework successfully reduces feature dimensionality in health data analysis.
- It effectively uncovers and validates causal relationships among health features.
- This contributes to improved healthcare and public health strategies through advanced data analysis.
More Related Videos
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Causality in Epidemiology
Comparing the Survival Analysis of Two or More Groups
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Statistical Methods for Analyzing Epidemiological Data
Criteria for Causality: Bradford Hill Criteria - II