Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

19.0K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.0K
DNA Microarrays02:34

DNA Microarrays

21.7K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
21.7K
Next-generation Sequencing03:00

Next-generation Sequencing

100.1K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
100.1K
Sanger Sequencing01:57

Sanger Sequencing

776.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
776.7K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

19.2K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
19.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Non-Parametric Ancestry Adjustment for Polygenic Scores.

medRxiv : the preprint server for health sciences·2026
Same author

The landscape of genomic and socioeconomic variables in colorectal cancer patients based on genetic ancestry.

Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive Oncology·2026
Same author

Analysis of a deeply-phenotyped familial hypercholesterolemia cohort from Mexico shows a role for both rare and common alleles across known dyslipidemia genes and reveals structural variation in a novel locus.

Human genomics·2025
Same author

Benchmarking of germline copy number variant callers from whole genome sequencing data for clinical applications.

Bioinformatics advances·2025
Same author

Diverse ancestral representation improves genetic intolerance metrics.

Nature communications·2025
Same author

Session Introduction: Overcoming health disparities in precision medicine: Intersectional approaches in precision medicine.

Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing·2024

Related Experiment Video

Updated: Mar 9, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.8K

Using genotype array data to compare multi- and single-sample variant calls and improve variant call sets from deep

Suyash S Shringarpure1, Rasika A Mathias2,3, Ryan D Hernandez4,5,6

  • 1Departments of Genetics and Biomedical Data Science, Stanford University School of Medicine, Stanford, CA, USA.

Bioinformatics (Oxford, England)
|December 31, 2016
PubMed
Summary

We developed a Random Forest classifier using genotype array data to improve variant calling accuracy from next-generation sequencing (NGS) data, achieving high true positive rates and offering adjustable quality criteria for different variant frequencies.

More Related Videos

Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
11:02

Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing

Published on: October 18, 2013

20.0K
Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
09:33

Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens

Published on: August 25, 2023

1.8K

Related Experiment Videos

Last Updated: Mar 9, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.8K
Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
11:02

Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing

Published on: October 18, 2013

20.0K
Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
09:33

Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens

Published on: August 25, 2023

1.8K

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Next-generation sequencing (NGS) variant calling is prone to false positives from technical errors.
  • Distinguishing true variants from errors is crucial for accurate genomic analysis.
  • Existing methods often rely on public databases, which may not fully represent specific study populations.

Purpose of the Study:

  • To develop and validate a novel method for improving variant call accuracy using sample-specific genotype array data.
  • To train a Random Forest classifier to differentiate true positive variant calls from false positives.
  • To assess the performance of the classifier across different variant calling algorithms and allele frequencies.

Main Methods:

  • Utilized genotype array data from sequenced samples to train a Random Forest classifier.
  • Applied the classifier to variant calls from 642 African-ancestry genomes (CAAPA study) sequenced at 30X depth.
  • Evaluated classifier performance against single-sample (CASAVA) and multi-sample callers (Real Time Genomics, GATK UnifiedGenotyper).

Main Results:

  • Achieved high true positive rates (97.5%, 95%, 99%) at a 5% false positive rate for CASAVA, Real Time Genomics, and GATK UnifiedGenotyper, respectively.
  • Demonstrated superior computational validation of site calls compared to generic methods.
  • Showcased the ability to adjust quality criteria based on allele frequency, providing insights into call quality for rare and common variants.

Conclusions:

  • The developed method effectively enhances variant calling accuracy by leveraging sample-specific genotype data.
  • This approach offers a more robust and adaptable computational validation of variant calls.
  • The method provides valuable insights into sequencing data quality across different variant frequencies.