Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Determination of Expected Frequency01:08

Determination of Expected Frequency

2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.6K
RNA-seq03:21

RNA-seq

10.3K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The genomic impact of population connectivity and decline in Africa's elephants.

Nature communications·2026
Same author

Population discontinuity in the Paris Basin linked to evidence of the Neolithic decline.

Nature ecology & evolution·2026
Same author

Flexible read-aware genotype imputation from sequence using biobank sized reference panels.

Nature communications·2025
Same author

SVUPP: Pre-phasing long reads improves structural variant genotyping.

Bioinformatics (Oxford, England)·2025
Same author

Polygenic risk score for type 2 diabetes shows context-dependent effects across populations.

Nature communications·2025
Same author

The genetic diversity of Indonesian cattle has been shaped by multiple introductions and adaptive introgression.

Nature communications·2025

Related Experiment Video

Updated: Aug 27, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.2K

Estimation of site frequency spectra from low-coverage sequencing data using stochastic EM reduces overfitting,

Malthe Sebro Rasmussen1, Genís Garcia-Erill1, Thorfinn Sand Korneliussen2

  • 1Department of Biology, University of Copenhagen, 2200 København N, Denmark.

Genetics
|September 29, 2022
PubMed
Summary

We developed a new algorithm to accurately estimate the site frequency spectrum (SFS) from low-coverage sequencing data, reducing computational demands and improving population genetics inferences.

Keywords:
demographic historyexpectation–maximizationgenotype likelihoodslow-coverage datanext-generation sequencingsite frequency spectrum

More Related Videos

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.3K
Enhanced Reduced Representation Bisulfite Sequencing for Assessment of DNA Methylation at Base Pair Resolution
13:47

Enhanced Reduced Representation Bisulfite Sequencing for Assessment of DNA Methylation at Base Pair Resolution

Published on: February 24, 2015

25.7K

Related Experiment Videos

Last Updated: Aug 27, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.2K
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.3K
Enhanced Reduced Representation Bisulfite Sequencing for Assessment of DNA Methylation at Base Pair Resolution
13:47

Enhanced Reduced Representation Bisulfite Sequencing for Assessment of DNA Methylation at Base Pair Resolution

Published on: February 24, 2015

25.7K

Area of Science:

  • Population Genetics
  • Bioinformatics
  • Genomics

Background:

  • The site frequency spectrum (SFS) is crucial for understanding population genetics, including demographic history and selection.
  • Estimating SFS from low-coverage sequencing data is challenging due to genotype calling biases.
  • Existing methods for SFS estimation can be computationally intensive and prone to overfitting.

Purpose of the Study:

  • To present a novel stochastic expectation-maximization algorithm for SFS inference from next-generation sequencing (NGS) data.
  • To address the computational and overfitting challenges associated with current SFS estimation methods.
  • To improve the accuracy of downstream population genetics inferences.

Main Methods:

  • Developed a stochastic expectation-maximization algorithm for SFS inference.
  • Designed the algorithm to handle low-coverage sequencing data and minimize bias.
  • Focused on reducing computational demands and memory usage for genome-scale analyses.

Main Results:

  • The proposed algorithm significantly reduces runtime for SFS estimation.
  • The method achieves constant and minimal RAM usage, enabling scalable analysis.
  • The algorithm effectively reduces overfitting, leading to more reliable downstream inferences.

Conclusions:

  • The stochastic expectation-maximization algorithm provides an efficient and accurate solution for SFS inference from NGS data.
  • This method overcomes key limitations of existing approaches, particularly for large-scale genomic datasets.
  • The improved SFS estimation facilitates more robust inferences in population genetics research.