Related Experiment Video
Updated: Nov 22, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
Reducing noise and stutter in short tandem repeat loci with unique molecular identifiers
August E Woerner1, Sammed Mandape2, Jonathan L King2
1Center for Human Identification, University of North Texas Health Science Center, 3500 Camp Bowie Blvd., Fort Worth, TX 76107, USA; Department of Microbiology, Immunology and Genetics, University of North Texas Health Science Center, 3500 Camp Bowie Blvd., Fort Worth, TX 76107, USA.
Unique molecular identifiers (UMIs) correct PCR and sequencing errors in DNA. This new bioinformatics tool, strumi, improves short tandem repeat (STR) analysis accuracy by creating consensus reads from UMI-tagged molecules.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Massively parallel sequencing (MPS) is prone to errors during PCR.
- Unique molecular identifiers (UMIs) are used to tag DNA molecules before PCR to track and correct these errors.
- Short tandem repeats (STRs) analysis is particularly affected by PCR and sequencing errors.
Purpose of the Study:
- To introduce strumi, a novel bioinformatics pipeline for analyzing UMI-tagged STRs.
- To improve the accuracy of STR analysis by correcting PCR and sequencing errors.
- To enable more precise quantification and characterization of genetic variations.
Main Methods:
- Development of strumi, an alignment-free, machine learning-driven algorithm.
- Clustering of MPS reads into UMI families.
- Inference of consensus super-reads representing original DNA molecules.
- Application of both threshold-based and machine learning approaches for accuracy estimation.
Main Results:
- Naïve threshold-based approaches yielded accurate super-reads (∼97% haplotype accuracy).
- Machine learning approaches further increased accuracy to ∼99.5%.
- UMIs significantly improve STR analysis accuracy compared to traditional MPS (∼78% without UMIs).
Conclusions:
- UMIs and the strumi pipeline can simplify probabilistic genotyping and reduce uncertainty in STR analysis.
- The high sensitivity of UMI-based methods allows for the characterization of contamination and somatic variations.
- New challenges arise in interpreting trace-level variations, including somatic stutter and contamination.

