Related Experiment Video
Updated: Feb 4, 2026

10:50
Directed Evolution Method in Saccharomyces cerevisiae: Mutant Library Creation and Screening
Published on: April 1, 2016
11.4K
Sc-ncDNAPred: A Sequence-Based Predictor for Identifying Non-coding DNA in Saccharomyces cerevisiae
Wenying He1, Ying Ju2, Xiangxiang Zeng2
1School of Computer Science and Technology, Tianjin University, Tianjin, China.
Frontiers in Microbiology
|September 28, 2018
Summary
Researchers developed Sc-ncDNAPred, a machine learning tool to accurately identify non-coding DNA (ncDNA) sequences in yeast genomes. This advancement supports synthetic biology by providing essential genomic data for DNA assembly and artificial life research.
Area of Science:
- Genomics
- Synthetic Biology
- Bioinformatics
Background:
- Genomics research is shifting towards genome synthesis, driven by advances in sequencing and synthetic biology.
- DNA assembly technology is crucial for artificial life but requires accurate identification of non-coding DNA (ncDNA) sequences.
- Experimental methods for ncDNA detection are costly for genome-wide analysis, necessitating computational approaches.
Purpose of the Study:
- To develop a machine learning-based method for accurate prediction of non-coding DNA sequences.
- To create a computational tool that supports the data requirements of DNA assembly technologies.
- To provide an accessible resource for researchers studying yeast genomics.
Main Methods:
- Collected a benchmark dataset of ncDNA sequences from *Saccharomyces cerevisiae* (yeast).
- Employed a support vector machine (SVM) learning method for sequence prediction.
- Evaluated feature extraction strategies including mono-, di-, tri-, tetra-, penta-, and hexamers to optimize predictor performance.
Main Results:
- Developed Sc-ncDNAPred, a predictor for *Saccharomyces cerevisiae* ncDNA sequences.
- Achieved a high prediction accuracy of 0.98.
- Identified an optimal feature extraction strategy for ncDNA prediction using SVM.
Conclusions:
- Sc-ncDNAPred offers an accurate and efficient computational method for identifying ncDNA sequences.
- The tool facilitates advancements in synthetic biology, particularly in DNA assembly and artificial life.
- An online web server is available for user convenience at http://server.malab.cn/Sc_ncDNAPred/index.jsp.
More Related Videos
Related Concept Videos
lncRNA - Long Non-coding RNAs
10.0K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
10.0K
lncRNA - Long Non-coding RNAs
3.7K
3.7K
DNA Base Pairing
33.3K
Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
33.3K
DNA Base Pairing
32.6K
32.6K
Cis-regulatory Sequences
11.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.8K
From DNA to Protein
22.4K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
22.4K

