Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Words in DNA sequences: some case studies based on their frequency statistics.

Srabashi Basu1, Debi Prosad Burma, Probal Chaudhuri

  • 1Theoretical Statistics and Mathematics unit, Indian Statistical Institute, 203 B.T. Road, Calcutta 700108, India. srabashi@isical.ac.in

Journal of Mathematical Biology
|June 5, 2003
PubMed
Summary

Analyzing DNA sequences using word frequencies offers an effective method for summarizing large genomic data. This approach aids in identifying genomic structures and evolutionary relationships, overcoming limitations of traditional sequence alignment methods.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Performance assessment of genomic island prediction tools with an improved version of Design-Island.

Computational biology and chemistry·2022
Same author

Prevalence, awareness, and control of hypertension in the slums of Kolkata.

Indian heart journal·2016
Same author

Classification based on hybridization of parametric and nonparametric classifiers.

IEEE transactions on pattern analysis and machine intelligence·2009
Same author

On detection and assessment of statistical significance of Genomic Islands.

BMC genomics·2008
Same author

Multiscale classification using nearest neighbor density estimates.

IEEE transactions on systems, man, and cybernetics. Part B, Cybernetics : a publication of the IEEE Systems, Man, and Cybernetics Society·2006
Same author

On visualization and aggregation of nearest neighbor classifiers.

IEEE transactions on pattern analysis and machine intelligence·2005

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Traditional DNA sequence analysis methods like alignment struggle with large, variable-sized genomes.
  • Effective statistical summarization is crucial for analyzing extensive DNA sequences.

Purpose of the Study:

  • To explore the utility of DNA word frequencies for analyzing large genomic sequences.
  • To detect genomic structural signatures and infer phylogenetic relationships using word distribution patterns.

Main Methods:

  • Statistical analysis of DNA sequences based on word frequencies.
  • Application to complete genomes (baker's yeast), ribosomal RNA sequences (prokaryotic and eukaryotic), and bacteriophage genomes.

Main Results:

Related Experiment Videos

  • DNA word frequencies effectively reduce dimensionality of large sequences.
  • Structural genomic information with biological significance is retained.
  • Variations in word distributions reflect phylogenetic relationships.
  • Conclusions:

    • DNA word frequency analysis is a valuable tool for large-scale genomic data.
    • This method overcomes limitations of sequence alignment for diverse and large datasets.
    • Potential for identifying novel biological insights and addressing statistical challenges in DNA word analysis.