Related Experiment Video
Updated: Aug 1, 2025

14:26
Genome-wide Purification of Extrachromosomal Circular DNA from Eukaryotic Cells
Published on: April 4, 2016
25.3K
Short human eccDNAs are predictable from sequences
Kai-Li Chang1, Jia-Hong Chen1,2, Tzu-Chieh Lin1
1Institute of Information Science, Academia Sinica, Taipei, 115, Taiwan.
Briefings in Bioinformatics
|April 24, 2023
Summary
Short extrachromosomal circular DNAs (eccDNAs) formation is not random. Deep learning models reveal predictable DNA sequence features, enabling accurate cross-dataset predictions and uncovering hidden similarities in genomic data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Short extrachromosomal circular DNAs (eccDNAs) are abundant in eukaryotic cells, but their formation mechanisms remain debated.
- Previous studies suggested random or near-random eccDNA origins, hindering biomarker development.
- Recent interest in eccDNAs has spurred research into their specific formation patterns.
Purpose of the Study:
- To investigate the predictability of short eccDNA formation using deep learning.
- To challenge the notion of random eccDNA origins by identifying underlying sequence features.
- To develop a computational framework for analyzing eccDNA predictability.
Main Methods:
- Developed DeepCircle, a bioinformatics framework utilizing convolution- and attention-based neural networks.
- Applied DeepCircle to analyze human eccDNA datasets.
- Trained models to predict eccDNA formation based on DNA sequence features.
Main Results:
- DeepCircle achieved high prediction accuracy for short human eccDNAs across datasets (convolutional models: 79.65±4.7%, attention-based models: 83.31±4.18%).
- Identified shared DNA sequence features predictive of eccDNA formation, despite low similarity in genomic locations.
- Demonstrated that eccDNA predictability is encoded within their sequences, irrespective of tissue origin.
Conclusions:
- The formation of short eccDNAs is intrinsically predictable and encoded in their DNA sequences.
- Deep learning models can uncover hidden similarities in genomic data, re-evaluating perceived lack of specificity.
- Findings support further investigation of eccDNAs as potential biomarkers.
Related Concept Videos
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
Gene Evolution - Fast or Slow?
7.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.2K
Evolutionary Relationships through Genome Comparisons
5.9K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.9K
Next-generation Sequencing
91.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
91.7K
Maxam-Gilbert Sequencing
11.3K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.3K
Organization of Genes
68.8K
Overview
68.8K

