Accurate prediction of potential druggable proteins based on genetic algorithm and Bagging-SVM ensemble classifier

Jianying Lin1, Hui Chen1, Shan Li2

  • 1College of Mathematics and Physics, Qingdao University of Science and Technology, Qingdao 266061, China; Artificial Intelligence and Biomedical Big Data Research Center, Qingdao University of Science and Technology, Qingdao 266061, China; Key Laboratory of Synthetic Biology, CAS Center for Excellence in Molecular Plant Sciences, Institute of Plant Physiology and Ecology, Shanghai Institutes for Biological Sciences, Chinese Academy of Sciences, Shanghai 200032, China.

Summary

Machine learning accurately predicts druggable proteins, accelerating drug discovery. This novel method combines feature extraction, genetic algorithms, and ensemble learning for efficient drug target identification.

Related Concept Videos

A Protocol for Computer-Based Protein Structure and Function Prediction16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Guidelines for computer based structural and functional characterization of protein using the I-TASSER pipeline is described. Starting from query protein sequence, 3D models are generated using multiple threading alignments and iterative structural assembly simulations. Functional inferences are thereafter drawn based on matches to proteins with known structure and...
69.7K
Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis08:49

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis

Computational methods hold promises for expediting drug discovery, yet they frequently overlook the dynamic nature of protein structures. Here, we discuss ensemble-based docking analysis to indirectly incorporate protein flexibility, potentially improving the accuracy and reliability of drug discovery...
1.1K
P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation06:09

P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation

This article presents a method for estimating same-day P300 speller Brain-Computer Interface (BCI) accuracy using a small testing dataset.
931
Classifying Matter by Composition03:35

Classifying Matter by Composition

Matter: Pure Substances and Mixtures
According to its composition, the matter can be classified into two broad categories — pure substances and mixtures. 
A pure substance is a form of matter that has a constant composition throughout with uniform properties. For example, any sample of sucrose has the same composition and same physical properties, such as melting point, color, and sweetness, regardless of the source from which it is isolated. 
A mixture is composed of two or...
89.5K
Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures07:05

Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures

We present a method, which utilizes a generalizable area-based image analysis approach to identify cell counts. Analysis of different cell populations exploited the significant cell height and structure differences between distinct cell types within an adaptive...
2.9K
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

OpenProt is a freely accessible database that enforces a polycistronic model of eukaryotic genomes. Here, we present a protocol for the use of OpenProt databases when interrogating mass spectrometry datasets. Using OpenProt database for analysis of proteomic experiments allows for discovery of novel and previously undetectable proteins.
13.3K