Related Experiment Video
Updated: Aug 14, 2025

07:51
High-throughput Identification of Synergistic Drug Combinations by the Overlap2 Method
Published on: May 21, 2018
11.9K
Multimodal data fusion for supervised learning-based identification of USP7 inhibitors: a systematic comparison
Wen-Feng Shen1, He-Wei Tang1, Jia-Bo Li1
1School of Medicine & School of Computer Engineering and Science, Shanghai University, Shanghai, 200444, China.
Journal of Cheminformatics
|January 11, 2023
Summary
This study developed accurate machine learning models to identify Ubiquitin-specific-processing protease 7 (USP7) inhibitors for cancer therapy. Ensemble and deep learning models showed promise, especially when handling imbalanced datasets, accelerating drug discovery.
Area of Science:
- Computational chemistry
- Drug discovery
- Machine learning
Background:
- Ubiquitin-specific-processing protease 7 (USP7) is a key target for cancer therapy.
- Developing USP7 inhibitors is challenging due to limitations in traditional virtual screening accuracy.
Purpose of the Study:
- To build accurate supervised learning models for identifying USP7 inhibitors.
- To evaluate different molecular representations and machine learning models for this task.
- To provide guidance on handling imbalanced datasets in drug discovery.
Main Methods:
- Calculated physicochemical descriptors, MACCS keys, ECFP4 fingerprints, and SMILES for compounds.
- Constructed two deep learning (DL) and nine classical machine learning (ML) models.
- Implemented 75 experiments across 15 groups using various molecular representations and activity cutoffs.
Main Results:
- Ensemble learning models performed optimally on balanced and severely imbalanced datasets.
- SMILES-based DL models excelled on slightly imbalanced datasets.
- Data fusion, SMOTE, unbiased decoy selection, and SMILES enumeration improved model performance, particularly for imbalanced data.
Conclusions:
- Highly accurate supervised learning models were established to accelerate USP7 inhibitor development.
- Guidance is provided for selecting models, molecular representations, and methods for imbalanced datasets in drug discovery research.
Keywords:
Deep learningMachine learningMolecular representationsMultimodal data fusionUbiquitin-specific-processing protease 7More Related Videos
Related Concept Videos
Tagging and Fusion Proteins
6.8K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.8K
Peptide Identification Using Tandem Mass Spectrometry
6.6K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.6K

