Related Experiment Video
Updated: Aug 6, 2025

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Improvement of multi-task learning by data enrichment: application for drug discovery
Ekaterina A Sosnina1, Sergey Sosnin2, Maxim V Fedorov3,4
1Center for Computational and Data-Intensive Science and Engineering, Skolkovo Institute of Science and Technology, Bolshoy Boulevard 30/1, Moscow, Russia, 143026. ekaterina.sosnina@skoltech.ru.
Training data enrichment can improve multi-task learning for drug discovery predictions. However, improvements depend on data quality and the inclusion of diverse compounds and targets for better accuracy.
Area of Science:
- Computational chemistry and cheminformatics
- Artificial intelligence in drug discovery
Background:
- Multi-task learning (MTL) is increasingly vital in drug discovery for predicting compound activities.
- Challenges remain in optimizing MTL models for enhanced prediction performance.
- Deep neural networks (DNNs) are powerful tools for complex predictive tasks in pharmacology.
Purpose of the Study:
- To investigate the impact of training data enrichment on multi-task deep neural network (MT-DNN) prediction quality in drug discovery.
- To evaluate how varying information capacity in training data affects model performance.
- To assess the generalizability of MT-DNNs for novel compound-target interactions.
Main Methods:
- Utilized three datasets: ViralChEMBL (classification), pQSAR(159), and pQSAR(4267) (regression).
- Developed feed-forward DNN-based multi-task models using the PyTorch framework.
- Evaluated four training data enrichment scenarios and two test data types to assess prediction accuracy.
Main Results:
- Training data enrichment is an effective strategy for enhancing MT-DNN prediction performance in drug discovery.
- The extent of performance improvement is contingent upon the quality and diversity (unique compounds and targets) of the training data.
- MT-DNNs struggle to predict interactions for compounds significantly dissimilar to those in the training set, even with enrichment.
Conclusions:
- Data enrichment strategies can significantly boost prediction accuracy in multi-task learning for drug discovery.
- Careful consideration of training data composition, including compound and target diversity, is crucial for maximizing benefits.
- Recommendations are provided for optimizing MT-DNNs to improve prediction accuracy and accelerate the identification of novel drug candidates.
Related Concept Videos
Drug Discovery: Overview
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Analysis of Population Pharmacokinetic Data
Drug Biotransformation: Overview
Targets for Drug Action: Overview
Receptors are either membrane-spanning or intracellular proteins, which upon binding a ligand, get activated and transmit the signal downstream to elicit a response. Drugs bind receptors, either mimicking the action of endogenous ligands or blocking the receptor activity to bring about a modified response. Nearly 35% of approved drugs target the G...
Time Course of Drug Effect

