Related Experiment Video
Updated: Feb 15, 2026

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins
Published on: August 9, 2019
Protein Sequence Comparison and DNA-binding Protein Identification with Generalized PseAAC and Graphical
Chun Li1,2,3, Jialing Zhao2, Changzhong Wang2
1School of Mathematics and Statistics, Hainan Normal University, Haikou 571158, China.
A new computational method, generalized pseudo amino acid composition (PseAAC), efficiently encodes protein sequences. This approach accurately compares protein similarities and identifies DNA-binding proteins, outperforming existing methods.
Area of Science:
- Bioinformatics
- Computational Biology
- Proteomics
Background:
- The exponential growth of protein sequence data necessitates advanced computational tools.
- Existing methods for protein sequence analysis face challenges in efficiency and accuracy.
Purpose of the Study:
- To develop an efficient computational approach for encoding protein sequences.
- To extract hidden information from protein sequences for analysis and comparison.
- To create a numerical descriptor for protein sequences.
Main Methods:
- Converted protein primary sequences into three-letter sequences based on physicochemical properties.
- Constructed a generalized pseudo amino acid composition (PseAAC) model.
- Developed a support vector machine (SVM) model using generalized PseAAC for DNA-binding protein identification.
Main Results:
- Successfully compared sequence similarities among beta-globin and coronavirus spike proteins, with clusters aligning to known taxonomic groups.
- The generalized PseAAC-based SVM model demonstrated superior performance in identifying DNA-binding proteins compared to four existing methods (DNAbinder, DNA-Prot, iDNA-Prot, enDNA-Prot).
- Achieved significant improvements in accuracy (ACC), Matthew's correlation coefficient (MCC), and F1-score (F1M) across various datasets.
Conclusions:
- The generalized PseAAC model is highly effective for protein sequence comparison and analysis.
- The developed method shows significant promise and competitiveness in the task of identifying DNA-binding proteins.
- This approach offers a valuable tool for bioinformatics research and applications.
Related Concept Videos
From DNA to Protein
Graphical Representation of Inequalities
Graphical and Analytic Representation of Sinusoids
The first step is measuring the peak-to-peak value, which is twice the amplitude of the sinusoid. This provides information about the maximum voltage swing of the waveform.
Secondly, the period and angular frequency are determined. The period is the time taken for one complete cycle of the waveform, while...
Single-Strand DNA Binding Proteins
Factors Affecting Protein-Drug Binding: Protein-Related Factors
The physicochemical properties of a drug play a significant role in its ability to bind to proteins. Lipophilic drugs, which dissolve in fats, oils, and lipids, can be...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

