Related Experiment Video
Updated: Aug 19, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.2K
Using unique molecular identifiers to improve allele calling in low-template mixtures
Benjamin Crysup1, Sammed Mandape1, Jonathan L King1
1Center for Human Identification, University of North Texas Health Science Center, 3500 Camp Bowie Blvd., Fort Worth, TX 76107, USA.
Forensic Science International. Genetics
|December 3, 2022
Summary
This study introduces a new bioinformatic pipeline using unique molecular identifiers (UMIs) and machine learning to reduce PCR artifacts in forensic DNA mixture analysis. The method significantly improves accuracy, especially for low-template samples.
Area of Science:
- Forensic Science
- Molecular Biology
- Bioinformatics
Background:
- Polymerase chain reaction (PCR) artifacts pose challenges in DNA sequencing, particularly for low-template samples and complex mixtures.
- These artifacts can hinder the accurate analysis of minor contributors in forensic samples.
- Molecular barcoding techniques, successful in medicine for detecting low-abundance somatic variation, offer potential for forensic applications.
Purpose of the Study:
- To develop and evaluate a bioinformatic pipeline for analyzing targeted short tandem repeat (STR) sequencing data from forensic DNA mixtures.
- To reduce the impact of PCR errors and artifacts in forensic mixture analysis.
- To assess the effectiveness of unique molecular identifiers (UMIs) and machine learning in improving data accuracy.
Main Methods:
- Targeted sequencing of selected short tandem repeat (STR) loci.
- Development of a bioinformatic pipeline incorporating unique molecular identifiers (UMIs).
- Application of machine learning algorithms to filter PCR noise and identify true alleles.
- Experimental validation using mixed DNA samples at various ratios and starting concentrations.
Main Results:
- The developed pipeline effectively utilizes UMI data to reduce noise and improve the analysis of PCR artifacts.
- Machine learning further enhanced the filtering of noise products, leading to a significant reduction in accepted noise alleles compared to raw UMI counts or read counts.
- Performance improvements were observed across different DNA mixture ratios and starting DNA amounts.
Conclusions:
- The integration of UMIs and machine learning provides a powerful approach to mitigate PCR artifacts in forensic DNA mixture analysis.
- This method enhances the reliability of sequencing data, particularly for challenging low-template and mixed samples.
- The pipeline demonstrates a significant improvement in distinguishing true alleles from noise, advancing forensic DNA analysis capabilities.

