Related Experiment Video
Updated: Jul 16, 2026

Detection of Homologous Recombination Intermediates via Proximity Ligation and Quantitative PCR in Saccharomyces cerevisiae
Published on: September 11, 2022
Discarding duplicate ditags in LongSAGE analysis may introduce significant error
Jeppe Emmersen1, Anna M Heidenblut, Annabeth Laursen Høgh
1Department of Biotechnology, Chemistry and Environmental Engineering, Aalborg University, Aalborg, Denmark. je@bio.aau.dk <je@bio.aau.dk>
Removing duplicate ditags in Serial Analysis of Gene Expression (SAGE) is often unnecessary and introduces significant errors in LongSAGE analysis. New algorithms can identify true artifact ditags for accurate gene expression profiling.
Area of Science:
- Molecular Biology
- Genomics
- Bioinformatics
Background:
- Serial Analysis of Gene Expression (SAGE) involves removing duplicate ditags, presumed artifacts, during analysis.
- This removal process can eliminate naturally occurring duplicate ditags, leading to measurement errors.
Purpose of the Study:
- To investigate the necessity of removing all duplicate ditags in SAGE and LongSAGE analysis.
- To develop and apply an algorithm for analyzing ditag populations and identifying artifact ditags.
Main Methods:
- Development of a novel algorithm to analyze the differential occurrence of SAGE tags within ditag combinations.
- Application of the algorithm to a pancreatic acinar cell LongSAGE library and ten additional LongSAGE libraries.
- Comparative analysis of error introduced by duplicate ditag removal in SAGE versus LongSAGE.
Main Results:
- Analysis revealed no general amplification bias justifying the removal of all duplicate ditags in the studied LongSAGE libraries.
- Removal of duplicate ditags introduced insignificant errors in SAGE but up to 3-fold errors in LongSAGE.
- The developed algorithm successfully identified artifact ditags originating from nucleotide variations and vector contamination.
Conclusions:
- The routine removal of all duplicate ditags is unfounded for the analyzed datasets and introduces substantial errors, particularly in LongSAGE.
- This practice may lead to similar errors in other existing LongSAGE datasets.
- Ditag population analysis is crucial for identifying and managing artifact tags, ensuring more accurate gene expression data.
Related Concept Videos
Detection of Gross Error: The Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...
Longitudinal Research
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are observed.
