Related Experiment Video
Updated: May 22, 2025

10:34
Ultra-long Read Sequencing for Whole Genomic DNA Analysis
Published on: March 15, 2019
22.6K
Pragmatic soft-decision data readout of encoded large DNA
Qi Ge1, Rui Qin1, Shuang Liu1
1School of Microelectronics, Tianjin University, No. 92 Weijin Road, Nankai District, Tianjin 300072, China.
Briefings in Bioinformatics
|March 17, 2025
Summary
This study introduces a novel DNA data storage readout method using soft-decision algorithms. It enables error-free DNA data recovery from noisy reads with ultra-low coverage, advancing DNA data archiving.
Area of Science:
- Biotechnology and Synthetic Biology
- Bioinformatics and Computational Biology
- Data Storage and Archiving
Background:
- DNA data storage offers high density and long-term stability for archiving.
- Nanopore sequencing provides rapid, long-read DNA sequencing but faces challenges with insertion/deletion (indel) errors and complex assembly.
- Efficient and accurate data readout from noisy, low-coverage DNA sequences remains a significant hurdle.
Purpose of the Study:
- To develop an assembly-free DNA data readout method for large DNA fragments.
- To correct insertion/deletion (indel) errors and enable ultra-low coverage data retrieval.
- To enhance the practicality and scalability of DNA data storage systems.
Main Methods:
- Embedding watermarks within large DNA fragments for direct read localization.
- Implementing a soft-decision forward-backward algorithm for indel error identification and correction.
- Utilizing minimum state transitions and read segmentation for rapid information extraction.
Main Results:
- Achieved assembly-free sequence reconstruction and error-free data recovery from noisy reads (~1% error rate) at 1-4× coverage for ~51 kb plasmids.
- Demonstrated robust performance and scalability through simulations on large-scale datasets under various error conditions.
- Enabled nearly single-molecule recovery, highlighting the method's efficiency.
Conclusions:
- The proposed soft-decision readout method overcomes limitations of traditional DNA data retrieval.
- This approach significantly improves the accuracy and efficiency of reading data from DNA storage.
- The method is particularly suitable for rapid readout applications in DNA data archiving.
Keywords:
DNA data storageencoded large DNAforward–backward algorithmhidden Markov modelsoft-decision strategyMore Related Videos
Related Concept Videos
Next-generation Sequencing
86.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
86.7K
DNA Isolation
37.5K
DNA isolation protocols can be fast and straightforward or complex and time-consuming depending on the type and quality of DNA required for further processing. For example, plasmid DNA extraction is a bit more complicated than genomic DNA extraction because of the need for an appropriate lysis method to separate plasmid DNA from gDNA during isolation. However, for specific applications, such as long-range DNA sequencing that require a good yield of high- quality DNA samples, we need to follow...
37.5K
Sanger Sequencing
752.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
752.0K
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K

