Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Genetic Variation01:25

Genetic Variation

1.1K
Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
1.1K
Next-generation Sequencing03:00

Next-generation Sequencing

97.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.4K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
15.2K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.4K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A blended genome and exome sequencing method captures genetic variation in an unbiased and cost-effective manner.

Nature genetics·2026
Same author

Thiazole-Linked <i>N</i>-Hydroxypropanamide Derivatives: Selective HDAC6 Inhibitors with Therapeutic Potential for Neurodegenerative Diseases.

Journal of medicinal chemistry·2026
Same author

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

Bioinformatics (Oxford, England)·2026
Same author

Effect of ancestry and shared genetic architecture of serious mental illness on symptoms and cognition in an admixed Latin American population.

medRxiv : the preprint server for health sciences·2026
Same author

A Global Prospective Harmonization Framework for Suicidality, Anhedonia, and Obsessive-Compulsive Symptoms in Psychiatric Genetic Studies: A Cross-Continental Study Within the Ancestral Population Network.

American journal of medical genetics. Part B, Neuropsychiatric genetics : the official publication of the International Society of Psychiatric Genetics·2026
Same author

Reward-Related Brain Function as an Endophenotype in the Mood-Psychosis Spectrum.

Biological psychiatry global open science·2026

Related Experiment Video

Updated: Jan 1, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

34.5K

ForestQC: Quality control on genetic variants from next-generation sequencing data using random forest.

Jiajin Li1, Brandon Jew2, Lingyu Zhan3

  • 1Department of Human Genetics, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA, United States of America.

Plos Computational Biology
|December 19, 2019
PubMed
Summary

ForestQC is a new tool that improves genetic variant quality control for next-generation sequencing (NGS) data. It uses machine learning to accurately identify and remove low-quality variants, enhancing the reliability of genetic studies.

More Related Videos

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.7K
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.5K

Related Experiment Videos

Last Updated: Jan 1, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

34.5K
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.7K
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.5K

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Next-generation sequencing (NGS) identifies most genetic variants but can produce low-quality data.
  • Poor quality variants in large genetic studies can lead to inaccurate findings.
  • Effective quality control is crucial for reliable genomic analysis.

Purpose of the Study:

  • To develop and evaluate ForestQC, a novel statistical tool for genetic variant quality control.
  • To combine traditional filtering with machine learning for improved variant quality assessment.
  • To enhance the accuracy and efficiency of variant filtering in large-scale sequencing projects.

Main Methods:

  • ForestQC utilizes sequencing quality metrics (depth, genotyping quality, GC content) to predict false-positive variants.
  • A hybrid approach combining traditional filtering and machine learning algorithms is employed.
  • The tool was validated on two whole-genome sequencing datasets (related and unrelated individuals).

Main Results:

  • ForestQC demonstrated superior performance compared to established methods like GATK's VQSR.
  • The tool significantly improved the quality of variants for downstream genetic analysis.
  • ForestQC is computationally efficient, suitable for large-scale sequencing datasets.

Conclusions:

  • Combining machine learning with filtering approaches provides a practical solution for genetic variant quality control.
  • ForestQC offers a robust and efficient method for enhancing variant data quality in NGS studies.
  • The tool contributes to more reliable and accurate genetic discoveries from sequencing data.