Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Estimating Population Standard Deviation01:26

Estimating Population Standard Deviation

3.5K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.5K
Random and Systematic Errors01:20

Random and Systematic Errors

926
926
Random and Systematic Errors01:20

Random and Systematic Errors

16.0K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
16.0K
Estimating Population Mean with Unknown Standard Deviation01:22

Estimating Population Mean with Unknown Standard Deviation

9.0K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
9.0K
Estimating Population Mean with Known Standard Deviation01:16

Estimating Population Mean with Known Standard Deviation

9.8K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
9.8K
Contaminants and Errors01:16

Contaminants and Errors

527
Effective sample preparation is crucial for accurate and reliable laboratory analysis. During this process, two significant sources of error can arise: concentration bias from improper sample splitting and contamination caused by methods used to reduce particle size, such as grinding or homogenization. Identifying and minimizing these potential errors is crucial to ensuring the validity of the analysis.
Another key consideration is determining the appropriate number of samples required to...
527

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Oridonin protects the lung against hyperoxia-induced injury in a mouse model.

Undersea & hyperbaric medicine : journal of the Undersea and Hyperbaric Medical Society, Inc·2017
Same author

Changes in angiotensin II and angiotensin-converting enzyme of different tissues after prolonged hyperoxia exposure.

Undersea & hyperbaric medicine : journal of the Undersea and Hyperbaric Medical Society, Inc·2017
Same author

Sarcomatoid Renal Cell Carcinoma Has a Distinct Molecular Pathogenesis, Driver Mutation Profile, and Transcriptional Landscape.

Clinical cancer research : an official journal of the American Association for Cancer Research·2017
Same author

Cell type-selective imaging and profiling of newly synthesized proteomes by using puromycin analogues.

Chemical communications (Cambridge, England)·2017
Same author

A cheat sheet to navigate the complex maze of pharmaceutical exclusivities in Europe.

Pharmaceutical patent analyst·2017
Same author

Hydroxamic Acids as Chemoselective (ortho-Amino)arylation Reagents via Sigmatropic Rearrangement.

Angewandte Chemie (International ed. in English)·2017

Related Experiment Video

Updated: Mar 22, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.7K

Empirical estimation of sequencing error rates using smoothing splines.

Xuan Zhu1, Jian Wang1, Bo Peng2

  • 1Department of Biostatistics, The University of Texas MD Anderson Cancer Center, Houston, TX, 77030, USA.

BMC Bioinformatics
|April 23, 2016
PubMed
Summary

We developed a new method to accurately estimate sequencing errors in next-generation sequencing data. This approach improves genomic analysis by not assuming a linear relationship between reads and errors.

Keywords:
Empirical error rateFrequency-based simulationNext-generation sequencingShort readsSmoothing spline

More Related Videos

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.8K
Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
09:30

Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms

Published on: September 13, 2018

10.0K

Related Experiment Videos

Last Updated: Mar 22, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.7K
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.8K
Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
09:30

Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms

Published on: September 13, 2018

10.0K

Area of Science:

  • Genomics
  • Bioinformatics

Background:

  • Next-generation sequencing (NGS) is crucial for biological discovery but suffers from higher error rates than conventional methods.
  • These errors complicate downstream genomic analyses, necessitating accurate error rate estimation.
  • Existing methods, like shadow regression, assume a linear relationship between reads and errors, which may not hold true for all data.

Purpose of the Study:

  • To develop a more reliable method for estimating error rates in next-generation sequencing data.
  • To overcome the limitations of linear assumptions in existing error estimation techniques.

Main Methods:

  • Proposed an empirical error rate estimation approach using cubic and robust smoothing splines.
  • Modeled the relationship between the number of sequenced reads and the number of reads containing errors (shadows).
  • Performed simulation studies using a frequency-based approach to generate realistic read and shadow counts.

Main Results:

  • The proposed approach demonstrated more accurate error rate estimations compared to shadow linear regression across all tested scenarios.
  • Simulation studies confirmed the robustness and accuracy of the empirical method.
  • Applied the approach to real-world datasets from major genomics projects (MAQC, ENCODE) and mutation screening studies.

Conclusions:

  • The novel empirical error rate estimation approach provides more accurate results for next-generation, short-read sequencing data.
  • This method does not rely on a linear relationship between error-free reads and shadow counts, offering broader applicability.
  • Improved error rate estimation enhances the reliability of genomic analyses derived from NGS data.