Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multi-species Conserved Sequences02:51

Multi-species Conserved Sequences

Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale  studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Sanger Sequencing01:57

Sanger Sequencing

DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
Next-generation Sequencing03:00

Next-generation Sequencing

The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Genome Size and the Evolution of New Genes03:21

Genome Size and the Evolution of New Genes

While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
DNA as a Genetic Template02:05

DNA as a Genetic Template

Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

CD4+ T cells in Type 1 Diabetes: Inferring Stage-specific Dysregulation from scRNA-seq.

Genomics, proteomics & bioinformatics·2026
Same author

Ancestral intronic splicing regulatory elements in the SCNα gene family.

RNA (New York, N.Y.)·2026
Same author

Inference of SARS-CoV-2 exposure biomarkers using large-scale T-cell repertoire profiling.

Genome medicine·2026
Same author

Ancestral intronic splicing regulatory elements in the SCN<i>α</i> gene family.

bioRxiv : the preprint server for biology·2026
Same author

Cardiomyopathies, heart rhythm and conduction disorders as phenotypic manifestation of genetic variants in large cohort of cardiac patients: results of whole-genome study.

Gene·2025
Same author

Interplay of the Genetic Variants and Allele Specific Methylation in the Context of a Single Human Genome Study.

International journal of molecular sciences·2025

Related Experiment Video

Updated: Jun 12, 2026

Novel Sequence Discovery by Subtractive Genomics
09:40

Novel Sequence Discovery by Subtractive Genomics

Published on: January 25, 2019

Exclusive sequences of different genomes.

Sergey I Mitrofanov1, Alexander Y Panchin, Sergei A Spirin

  • 1Faculty of Bioengineering and Bioinformatics, Moscow State University, Moscow, Russia. mitroser04@mail.ru

Journal of Bioinformatics and Computational Biology
|June 18, 2010
PubMed
Summary

We analyzed DNA sequences across many species to find over- and under-represented short DNA words. Our findings reveal consistent patterns and exceptions in genomic word usage across diverse life forms.

More Related Videos

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
06:40

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome

Published on: March 22, 2018

Optimization and Comparative Analysis of Plant Organellar DNA Enrichment Methods Suitable for Next-generation Sequencing
12:33

Optimization and Comparative Analysis of Plant Organellar DNA Enrichment Methods Suitable for Next-generation Sequencing

Published on: July 28, 2017

Related Experiment Videos

Last Updated: Jun 12, 2026

Novel Sequence Discovery by Subtractive Genomics
09:40

Novel Sequence Discovery by Subtractive Genomics

Published on: January 25, 2019

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
06:40

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome

Published on: March 22, 2018

Optimization and Comparative Analysis of Plant Organellar DNA Enrichment Methods Suitable for Next-generation Sequencing
12:33

Optimization and Comparative Analysis of Plant Organellar DNA Enrichment Methods Suitable for Next-generation Sequencing

Published on: July 28, 2017

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Understanding DNA sequence composition is crucial for deciphering genomic function.
  • Identifying non-random patterns in DNA word frequencies can reveal evolutionary constraints and biological significance.
  • Previous studies have explored DNA word distributions, but comprehensive analysis across diverse eukaryotic genomes is ongoing.

Purpose of the Study:

  • To investigate the distribution of short DNA words (1-7 base pairs) in a large dataset of eukaryotic genomes.
  • To identify statistically over- and under-represented DNA words using various modeling approaches.
  • To characterize the consistency and exceptions of these word patterns across different taxonomic groups.

Main Methods:

  • Analysis of a comprehensive dataset including 139 complete eukaryotic genomes, 33 masked genomes, and coding regions from 35 genomes.
  • Application and comparison of different statistical models for identifying over- and under-represented DNA words.
  • Utilizing the Karlin et al. method for its superior predictive power in detecting sequence biases.

Main Results:

  • The Karlin et al. method demonstrated the strongest predictive power for identifying biased DNA words.
  • Over- and under-represented words were identified with consistency across a wide range of taxonomic groups.
  • Novel over-represented words, such as CGCG in CG-deficient organisms, were discovered, alongside exceptions to widely recognized patterns like CG and TA.

Conclusions:

  • Statistical modeling of DNA word distribution provides valuable insights into genomic composition.
  • Consistent patterns of over- and under-represented words exist across eukaryotic genomes, suggesting underlying biological or evolutionary pressures.
  • The study highlights both conserved and novel sequence biases, including exceptions to previously established DNA word rules.