Related Experiment Video
Updated: Aug 12, 2025

08:57
Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
16.0K
Applying Machine Learning to Classify the Origins of Gene Duplications
Michael T W McKibben1, Michael S Barker2
1Department of Ecology & Evolutionary Biology, University of Arizona, Tucson, AZ, USA.
Methods in Molecular Biology (Clifton, N.J.)
|January 31, 2023
Summary
A new machine learning tool, Frackify, accurately classifies whole-genome duplication (WGD) paleologs in plant genomes. This method efficiently identifies ancient gene duplications, aiding evolutionary studies.
Area of Science:
- Genomics
- Evolutionary Biology
- Bioinformatics
Background:
- Whole-genome duplication (WGD) events are common in land plant evolution, leaving legacies in extant diploidized genomes.
- Genes originating from WGDs (paleologs) influence plant evolution through functional divergence, genetic diversity, and gene loss.
- Existing paleolog classification methods are limited by reliance on specific genomic features, accuracy, and computational demands.
Purpose of the Study:
- To develop a supervised machine learning approach for classifying paleologs from specific WGD events in diploidized plant genomes.
- To create a robust and efficient tool applicable across diverse paleopolyploidy histories, including overlapping WGDs.
Main Methods:
- Collected empirical genomic data (syntenic block sizes, etc.) from 27 plant species with varied paleopolyploidy histories.
- Developed simulations of syntenic blocks and paleologs to train a gradient boosted decision tree model.
- Implemented the Frackify (Fractionation Classify) tool for paleolog identification and classification.
Main Results:
- Frackify accurately identifies and classifies paleologs across a wide range of parameter spaces, including complex scenarios with multiple overlapping WGDs.
- Comparison with existing paleolog inference methods in six species demonstrated Frackify's high consistency and efficiency.
- The tool effectively combines multiple genomic features for rapid paleolog classification.
Conclusions:
- Frackify offers an accurate, efficient, and consistent method for classifying paleologs derived from whole-genome duplication events.
- This tool advances the study of paleologs and their impact on plant genome evolution.
- Frackify provides a valuable resource for researchers investigating ancient duplication events in plants.
Related Concept Videos
Gene Duplication and Divergence
6.2K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.2K
Gene Families
8.9K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.9K
Evolutionary Relationships through Genome Comparisons
6.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.1K
Gene Evolution - Fast or Slow?
7.3K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.3K
Exon Recombination
3.7K
The evolution of new genes is critical for speciation. Exon recombination, also known as exon shuffling or domain shuffling, is an important means of new gene formation. It is observed across vertebrates, invertebrates, and in some plants such as potatoes and sunflowers. During exon recombination, exons from the same or different genes recombine and produce new exon-intron combinations, which might evolve into new genes.
Exon shuffling follows “splice frame rules.” Each exon...
Exon shuffling follows “splice frame rules.” Each exon...
3.7K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K

