Related Experiment Video
Updated: Aug 15, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
dForml(KNN)-PseAAC: Detecting formylation sites from protein sequences using K-nearest neighbor algorithm via Chou's
Qiao Ning1, Zhiqiang Ma1, Xiaowei Zhao1
1Information Science and Technology, Northeast Normal University, Changchun 130117, China.
Journal of Theoretical Biology
|March 19, 2019
Summary
Researchers developed LFPred, a computational tool to predict lysine formylation sites in proteins. This method aids in understanding post-translational modifications crucial for biological functions.
Area of Science:
- Biochemistry
- Computational Biology
- Proteomics
Background:
- Post-translational modifications (PTMs) like protein formylation on lysine residues are vital for cellular functions.
- Accurate identification of formylation sites is essential for understanding these biological mechanisms.
- Existing experimental methods for site identification are time-consuming and less efficient than computational approaches.
Purpose of the Study:
- To develop the first computational predictor, LFPred, for identifying lysine formylation sites in proteins.
- To provide a faster and more convenient alternative to experimental methods for formylation site prediction.
Main Methods:
- Utilized sequence features including amino acid composition (AAC), binary profile features (BPF), and amino acid index (AAI).
- Employed the K-nearest neighbor (KNN) algorithm as the classification model.
- Implemented a discrete window approach based on information entropy and addressed data imbalance with carefully selected negative samples.
Main Results:
- The LFPred predictor achieved a specificity of 79.9% and a sensitivity of 81.4% in jackknife validation.
- Demonstrated the effectiveness of combining sequence features and the KNN algorithm for lysine formylation prediction.
Conclusions:
- LFPred serves as a valuable computational tool for accurate prediction of lysine formylation sites.
- The developed method offers a significant advancement in the study of protein formylation and its biological implications.
More Related Videos
Related Concept Videos
Protein Folding
Overview
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...

