Positive and negative forms of replicability in gene network analysis
W Verleyen1, S Ballouz1, J Gillis1
1Stanley Institute for Cognitive Genomics, Cold Spring Harbor Laboratory, 500 Sunnyside Boulevard Woodbury, NY 11797, USA.
Bioinformatics (Oxford, England)
|December 16, 2015
Summary
Research into gene networks may lead to overfitting. Targeting replication directly can result in uninformative findings, with replicability sometimes negatively correlating with accuracy in gene function prediction.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Gene networks are crucial for genomic data analysis but challenging to interpret.
- Extensive comparative evaluations and research into best practices may inadvertently lead to overfitting in the field.
Purpose of the Study:
- To investigate the potential for overfitting in gene network analysis.
- To model 'research communities' to understand performance trends in gene network analysis and replication.
Main Methods:
- Construction of a 'research community' model using real gene network data and machine learning.
- Analysis of performance trends related to replication and accuracy.
- Examination of factors influencing replicability in prior gene network studies.
Main Results:
- Directly targeting replication can lead to the dominance of uninformative findings.
- The relationship between replicability and accuracy is positive when network data and algorithms have similar variability (rs ≈ 0.33).
- Without such constraints, the relationship between replicability and accuracy can become negative for specific gene functions (rs ≈ -0.13).
- Factors driving replicability in previous studies are often linked to biases rather than correctness.
- Highly replicable interactions in protein-protein interaction data are frequently associated with poor quality control.
Conclusions:
- Overfitting is a potential issue in gene network research due to a focus on replication.
- Replicability does not always equate to accuracy and can be influenced by biases.
- Careful consideration of data quality and analytical methods is essential for reliable gene network analysis.
More Related Videos
Related Concept Videos
Protein Networks
4.7K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.7K
Protein Networks
2.9K
2.9K
Epistasis Analysis
6.2K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
6.2K
Genetic Screens
5.9K
Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
5.9K
Comparing Copy Number Variations and SNPs
19.2K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.2K
Gene Duplication and Divergence
8.2K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
8.2K


