Related Experiment Video
Updated: Jun 14, 2026

09:40
Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Assembly free comparative genomics of short-read sequence data discovers the needles in the haystack
Charles H Cannon1, Chai-Shian Kua, D Zhang
1Chinese Academy of Sciences, Menglun, Mengla, China. chuck@xbtg.ac.cn
Molecular Ecology
|March 25, 2010
Summary
This study introduces an assembly-free method for analyzing short-read sequence (SRS) data, identifying genomic variants without needing a reference genome. This approach enhances comparative genomics for non-model organisms and reveals population-level genetic markers.
Area of Science:
- Genomics
- Bioinformatics
- Population Genetics
Background:
- Comparative genomic analyses traditionally require a pre-assembled reference sequence.
- Short-read sequence (SRS) data is widely used but often limited by reliance on reference genomes.
- Analyzing genomic diversity in non-model organisms presents significant challenges.
Purpose of the Study:
- To develop and present an assembly-free method for analyzing SRS data.
- To discover sequence variants and compare genomic diversity across populations and families.
- To enhance the application of SRS technology for non-model organisms.
Main Methods:
- Developed an assembly-free analysis of SRS data by tabulating the presence and frequency of 'complex' fragments.
- Applied the method to SRS data from nine tree species, comparing population and family genomic diversity.
- Utilized simulated SRS data from three known plant genomes as a control.
Main Results:
- Identified three types of informative complexmers (Type I, II, III) with distinct statistical properties for variant detection.
- Type II complexmers highlighted potential copy-number differences between genomes.
- Type III complexmers proved useful for associating genetic differences with phenotypic or geographic variation, identifying markers for geographic origin in an endangered timber species.
Conclusions:
- The assembly-free approach significantly enhances SRS data utility for non-model organisms.
- The method directly identifies informative genetic elements for further study and assembly.
- Genomic divergence patterns observed in fig and stone oak species suggest ecological influences on gene flow and diversity.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...

