Related Experiment Video
Updated: Jul 15, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.1K
A digital twin for DNA data storage based on comprehensive quantification of errors and biases
Andreas L Gimpel1, Wendelin J Stark1, Reinhard Heckel2
1Department of Chemistry and Applied Biosciences, ETH Zürich, Vladimir-Prelog-Weg 1-5, 8093, Zürich, Switzerland.
Nature Communications
|September 27, 2023
Summary
This study characterizes errors in synthetic DNA data storage and creates a digital twin to simulate these processes. This enables data-driven development of error correction coding (ECC) and reduces costs.
Area of Science:
- Biotechnology
- Bioinformatics
- Data Storage
Background:
- Synthetic DNA offers high-density, long-term data archiving.
- Errors and biases in DNA data storage necessitate error correction coding (ECC), increasing redundancy.
- Lack of error data and modeling tools hinders ECC development.
Purpose of the Study:
- To comprehensively characterize error sources and biases in DNA data storage workflows.
- To develop a digital twin for simulating DNA data storage processes.
- To enable data-driven ECC development and optimize redundancy strategies.
Main Methods:
- Characterization of errors from DNA synthesis, PCR, aging, and sequencing.
- Development of a digital twin using data from 40 sequencing experiments.
- Simulation of common DNA data storage workflows and experimental reproduction.
Main Results:
- Identified and quantified errors and biases across key DNA data storage steps.
- Validated the digital twin's ability to accurately simulate experimental workflows.
- Demonstrated the digital twin's utility in replacing experiments and informing ECC design.
Conclusions:
- A digital twin can accurately model DNA data storage workflows, reducing the need for physical experiments.
- This approach facilitates data-driven development of more efficient error correction coding (ECC).
- Optimized ECC design leads to tangible cost savings in DNA data storage systems.
Related Concept Videos
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
DNA as a Genetic Template
22.0K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
22.0K
Next-generation Sequencing
89.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
89.8K
Proofreading
54.2K
Overview
54.2K

