Measuring the wisdom of the crowds in network-based gene function inference.
W Verleyen1, S Ballouz1, J Gillis1
1Stanley Institute for Cognitive Genomics, Cold Spring Harbor Laboratory, Woodbury, NY 11797, USA.
Bioinformatics (Oxford, England)
|November 1, 2014
Summary
Gene function prediction methods show similar performance when data is controlled. Aggregating data, not algorithms, significantly improves results, suggesting limited value in developing new prediction algorithms.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Network-based gene function inference methods are widely used but progress is unclear.
- Controlling for data and algorithm implementation is crucial for performance evaluation.
Purpose of the Study:
- To evaluate the performance trends of gene function inference methods.
- To assess the impact of data and algorithm aggregation on prediction accuracy.
Main Methods:
- Utilized well-characterized algorithms to generate 'untweaked' results.
- Controlled for underlying biological network data across different tests.
- Measured performance using standard metrics like area under the ROC curve (AUROC).
Main Results:
- State-of-the-art machine learning methods achieve 'gold standard' performance.
- Algorithm performance is highly aligned when controlling for data.
- Algorithm aggregation offers modest benefits (17% AUROC increase).
- Data aggregation yields substantial gains (88% AUROC improvement).
Conclusions:
- Additional algorithm development offers little improvement for gene function prediction.
- Data aggregation is a more effective strategy for enhancing prediction accuracy.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
12.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.1K
Protein Networks
3.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.6K
Protein Networks
1.8K
1.8K
Comparing Copy Number Variations and SNPs
11.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.4K
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Gene Evolution - Fast or Slow?
6.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
6.2K


