Related Experiment Video
Updated: Apr 14, 2026

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency
Published on: June 16, 2011
ViVaMBC: estimating viral sequence variation in complex populations from illumina deep-sequencing data using
Bie Verbist1, Lieven Clement2, Joke Reumers3
1Department of Mathematical Modeling, Statistics and Bioinformatics, Ghent University, Coupure Links 653, Gent, 9000, Belgium. Bie.Verbist@ugent.be.
ViVaMBC accurately identifies low-frequency viral variants using quality scores and codon-level analysis. This method enhances viral variant detection sensitivity, overcoming limitations in deep sequencing data.
Area of Science:
- Genomics
- Bioinformatics
- Virology
Background:
- Deep-sequencing technologies enable comprehensive analysis of sequence variations in populations.
- Sequencing errors can hinder accurate detection of low-frequency mutations.
- Illumina sequencing quality scores and secondary base calls offer potential for error correction.
Purpose of the Study:
- To introduce ViVaMBC, a novel model-based clustering method for identifying and quantifying viral variants.
- To leverage quality scores and second-best base calls for enhanced variant detection.
- To enable codon-level variant calling for direct biological interpretation, particularly for antiviral drug responses.
Main Methods:
- Developed ViVaMBC, a virus variant model-based clustering approach.
- Utilized Illumina sequencing quality scores and second-best base calls.
- Optimized variant calling at the codon level for biological relevance.
Main Results:
- ViVaMBC accurately estimates variant frequencies down to 0.5% with unbiased results at 25,000x coverage.
- Demonstrated superior sensitivity and specificity compared to V-Phaser2, ShoRAH, and LoFreq for variants >0.4%.
- Observed increased false positives below 0.4%, potentially due to sample/library preparation errors.
Conclusions:
- ViVaMBC is the first method for direct, codon-level viral variant calling.
- Quality score-based error modeling is the primary strength, significantly reducing sequencing errors.
- Second-best base calls offered marginal sensitivity gains, not justifying computational overhead; PCR errors become a limiting factor.
More Related Videos
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Sanger Sequencing
Size and Structure of Viral Genomes
Genetic Variation
Genes exist in different versions called alleles,...

