Related Experiment Video
Updated: Jul 30, 2025

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.9K
CurSa: scripts to curate metadata and sample genomes from GISAID for analysis and display in nextstrain and
1Department of Genetic Engineering, CINVESTAV-Irapuato, Irapuato, Guanajuato 36824, Mexico.
Biology Methods & Protocols
|May 14, 2023
Summary
This study introduces CurSa, a Perl script suite to clean and sample SARS-CoV-2 genome data. It addresses metadata errors, aiding accurate phylogenetic and evolutionary studies of the virus.
Area of Science:
- Genomics
- Bioinformatics
- Epidemiology
Background:
- SARS-CoV-2 has millions of sequenced genomes, posing bioinformatic challenges.
- Accurate geographic metadata is crucial for SARS-CoV-2 evolutionary and phylogenetic studies.
- Manual metadata entry into GISAID introduces errors, hindering research.
Purpose of the Study:
- To provide a tool for curating geographic information in SARS-CoV-2 sequence metadata.
- To facilitate random sampling of genome sequences for specific countries.
- To streamline data preparation for Nextstrain and Microreact, accelerating evolutionary analyses.
Main Methods:
- Development of a Perl script suite named CurSa.
- Implementation of functions for geographic data curation and sequence sampling.
- Designed for integration with Nextstrain and Microreact workflows.
Main Results:
- CurSa enables efficient correction of location-based metadata errors in SARS-CoV-2 genomic data.
- The scripts allow for targeted sampling of sequences from specific countries.
- Facilitates the preparation of high-quality datasets for phylogenetic analysis.
Conclusions:
- CurSa simplifies the curation of SARS-CoV-2 genomic metadata, improving data accuracy.
- The tool accelerates evolutionary studies by providing clean, well-sampled datasets.
- Enhances the reliability of SARS-CoV-2 phylogenetic and geographic analyses.
Related Concept Videos
Genome Annotation and Assembly
19.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.0K
Next-generation Sequencing
91.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
91.7K

