Related Experiment Video
Updated: Dec 28, 2025

Mapping Bacterial Functional Networks and Pathways in Escherichia Coli using Synthetic Genetic Arrays
Published on: November 12, 2012
Gene-gene interaction: the curse of dimensionality
Amrita Chattopadhyay1, Tzu-Pin Lu1
1Institute of Epidemiology and Preventive Medicine, Department of Public Health, National Taiwan University, Taipei.
Addressing the "missing heritability" in complex diseases requires analyzing gene-gene interactions (epistasis). PySpark and parallel computing offer solutions to overcome the curse of dimensionality in genome-wide association studies.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Genome-wide association studies (GWAS) often identify genetic variants with modest effects, contributing to the
- missing heritability
- problem in complex diseases.
- Gene-gene interactions (epistasis) are crucial for understanding complex disease etiology and identifying potential drug targets.
- Analyzing millions of single nucleotide polymorphisms (SNPs) for epistasis faces the
- curse of dimensionality
- , overwhelming traditional statistical methods.
Purpose of the Study:
- To explore advanced computational methods for analyzing gene-gene interactions and mitigating the challenges of high-dimensional genomic data.
- To highlight the potential of parallel computing and specific tools like PySpark in addressing the limitations of existing epistasis analysis methods.
Main Methods:
- Review of existing methods for epistasis analysis, including multifactor dimensionality reduction (MDR) and machine learning (ML) approaches (e.g., random forests, neural networks, deep learning).
- Discussion of the limitations of exhaustive search and variable selection in MDR and ML methods, and the interpretability issues with deep learning.
- Introduction of PySpark and distributed computing as a strategy to handle large datasets and improve processing speed for whole-genome gene-gene interaction studies.
Main Results:
- Traditional and MDR-based methods, as well as ML techniques, face challenges in exhaustive SNP interaction analysis due to high dimensionality.
- Deep learning methods, while powerful, present interpretability challenges.
- PySpark, leveraging distributed computing, offers a scalable solution to process vast amounts of genomic data efficiently, reducing information loss.
Conclusions:
- Overcoming the
- curse of dimensionality
- in epistasis analysis is critical for understanding complex diseases.
- Parallel computing frameworks like PySpark provide a robust approach to analyze large-scale genomic data, enabling more effective identification of gene-gene interactions.
- These advancements are essential for discovering novel gene functions, pathways, and therapeutic targets.
More Related Videos
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
09:23Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
Related Concept Videos
Epistasis Analysis
Gene-Environment Interactions
Polygenic Traits
Combinatorial Gene Control
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
Structure of a Gene
However, only 1% of the DNA is composed of genes that encode proteins; the rest, 99% is non-coding DNA. This non-coding DNA performs...
Position-effect Variegation