Related Experiment Video
Updated: Apr 23, 2026

08:57
Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
14.3K
An improved alignment-free model for DNA sequence similarity metric.
Junpeng Bao1, Ruiyu Yuan, Zhe Bao
1Department of Computer Science and Technology Xi'an Jiaotong University, West Xianning Road, 710049 Xi'an, P,R, China. baojp@mail.xjtu.edu.cn.
BMC Bioinformatics
|September 29, 2014
Summary
A new Category-Position-Frequency (CPF) model enhances DNA clustering by integrating nucleotide classification, position, and frequency. This alignment-free approach improves sequence similarity metrics for better data mining results.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- DNA clustering is crucial for analyzing large-scale DNA sequences.
- Current alignment-free methods often overlook valuable sequence information.
- Improving DNA sequence similarity metrics is key to enhancing clustering quality.
Purpose of the Study:
- To introduce a novel alignment-free model for DNA clustering.
- To enhance DNA sequence similarity metrics by incorporating compound information.
- To improve the overall quality of DNA clustering.
Main Methods:
- Proposed the Category-Position-Frequency (CPF) model.
- Converted DNA sequences into three category-based sequences.
- Generated a 12-dimension feature vector using an entropy-based model considering word frequency and position.
Main Results:
- The CPF model demonstrated superior performance compared to mainstream alignment-free models (k-tuple, DMk, TSM, AMI, CV).
- Achieved better clustering results and identified optimal settings.
- Validated through experiments on multiple DNA sequence datasets.
Conclusions:
- Hybrid information models outperform frequency-only models for DNA clustering.
- A sliding window size of two is optimal for the CPF model with sequences up to 5000 characters.
- The CPF model offers efficient, stable performance and broad generalization capabilities.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Modern Molecular Taxonomy
829
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
829
DNA as a Genetic Template
22.0K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
22.0K
RNA-seq
9.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.2K
Sanger Sequencing
800.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
800.2K

