Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Compression of nucleic acid and protein sequence data.

J R Walker1, P Willett

  • 1Department of Information Studies, University of Sheffield, Western Bank, UK.

Computer Applications in the Biosciences : CABIOS
|June 1, 1986
PubMed
Summary

This study applied text compression techniques, including n-gram and run-length coding, to biological sequence data. Combined methods achieved a 74.6% reduction in GenBank database size, significantly improving storage efficiency.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Animal movement on the hoof and on the cart and its implications for understanding exchange within the Indus Civilisation.

Scientific reports·2024
Same author

Dental Nutrition.

The Independent practitioner·2023
Same author

Quickest Detection of COVID-19 Pandemic Onset.

IEEE signal processing letters·2021
Same author

Reply to Rubber Company's Chemical Advocate.

The American journal of dental science·2019
Same author

On Emetics in Diabetes Mellitus.

The Medical and physical journal·2018
Same author

Predictors of patient reluctance to wake early in the morning for bowel preparation for colonoscopy: a precolonoscopy survey in city-wide practice.

Endoscopy international open·2018

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Data Science

Background:

  • Biological sequence data (nucleic acid and protein) is rapidly growing.
  • Efficient storage and management of large sequence databases are critical.
  • Current storage methods may not be optimal for the scale of genomic and proteomic data.

Purpose of the Study:

  • To evaluate the effectiveness of text compression methods for biological sequence data.
  • To reduce the storage requirements of machine-readable sequence files.
  • To assess the compression ratios achieved by specific algorithms.

Main Methods:

  • Application of n-gram coding for sequence data compression.
  • Implementation of run-length coding for sequence data compression.
  • Development of a Pascal program integrating both n-gram and run-length coding.

Main Results:

  • A combined n-gram and run-length coding approach achieved a 74.6% compression ratio on the GenBank database.
  • N-gram coding alone resulted in a 42.8% compression ratio for the Protein Identification Resource database.
  • Demonstrated significant data size reduction for biological sequence archives.

Conclusions:

  • Text compression techniques, particularly combined n-gram and run-length coding, are highly effective for reducing biological sequence data storage.
  • These methods offer practical solutions for managing large-scale genomic and proteomic databases.
  • Optimized data compression is essential for efficient bioinformatics research and data sharing.

Related Experiment Videos