Related Experiment Video
Updated: Apr 28, 2026

08:03
Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
2.1K
Using association rule mining to determine promising secondary phenotyping hypotheses
Anika Oellrich1, Julius Jacobsen1, Irene Papatheodorou1
1Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton CB1 10SA, UK.
Bioinformatics (Oxford, England)
|June 17, 2014
Summary
This study introduces an association rule mining method to predict novel gene-phenotype relationships. The approach identifies 1967 secondary phenotype hypotheses, offering valuable candidates for experimental validation in genetic research.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Large-scale phenotyping projects aim to link genes to phenotypes for understanding human diseases and drug development.
- Identifying gene-phenotype relations is resource-intensive due to the vast number of genes in vertebrate genomes.
Purpose of the Study:
- To develop a computational method for identifying promising secondary phenotype candidates for experimental testing.
- To leverage existing gene-phenotype annotations to discover novel associations.
Main Methods:
- Association rule mining applied to a comprehensive gene-phenotype annotation dataset.
- Identification of patterns in phenotype occurrences to predict new relationships.
- Evaluation of predicted candidates using automated and manual strategies.
Main Results:
- 1967 secondary phenotype hypotheses were generated, linking 244 genes and 136 phenotypes.
- The predicted secondary phenotypes demonstrated biological relevance to their associated genes.
- The findings suggest these predictions are suitable for experimental validation.
Conclusions:
- Association rule mining is an effective approach for predicting gene-phenotype relationships.
- The identified secondary phenotypes serve as strong candidates for future experimental research.
- This method aids in prioritizing experimental efforts in large-scale genetic studies.
More Related Videos
Related Concept Videos
Multiple Allele Traits
32.5K
The Concept of Multiple Allelism
32.5K
Epistasis Analysis
4.9K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
4.9K
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Epistasis
37.2K
In addition to multiple alleles at the same locus influencing traits, numerous genes or alleles at different locations may interact and influence phenotypes in a phenomenon called epistasis. For example, rabbit fur can be black or brown depending on whether the animal is homozygous dominant or heterozygous at a TYRP1 locus. However, if the rabbit is also homozygous recessive at a locus on the tyrosinase gene (TYR), it will have an unshaded coat that appears white, regardless of its TYRP1...
37.2K
Probability Laws
29.7K
Overview
29.7K
Hardy-Weinberg Principle
62.4K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
62.4K

