Related Experiment Video
Updated: Jan 8, 2026

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Small data, big challenges: Machine- and deep-learning strategies for data-limited drug discovery
Nazreen Pallikkavaliyaveetil1, Sriram Chandrasekaran2
1Department of Biomedical Engineering, University of Michigan, Ann Arbor, MI 48109, USA; Michigan Institute for Data & AI in Society (MIDAS), University of Michigan, Ann Arbor, MI 48109, USA.
The scarcity of high-quality experimental data, known as the small data problem, limits Machine Learning (ML) and Deep Learning (DL) in drug discovery. This review explores ML and DL strategies to overcome data limitations in drug discovery and development.
Area of Science:
- Computational chemistry
- Pharmacology
- Bioinformatics
Background:
- The drug discovery and development (DDD) pipeline is significantly hampered by a scarcity of high-quality experimental data, a common issue due to high costs, time, and confidentiality.
- Standard Machine Learning (ML) and Deep Learning (DL) algorithms struggle with small datasets, leading to challenges in feature extraction and model generalization.
- Existing reviews on AI in drug discovery often overlook the specific challenges posed by limited data across the entire DDD pipeline.
Purpose of the Study:
- To address the critical challenge of limited data in the drug discovery and development (DDD) pipeline.
- To survey key drug discovery tasks where data scarcity is prevalent.
- To synthesize both traditional ML and advanced DL strategies specifically adapted for small datasets in DDD.
Main Methods:
- Reviewing existing literature on AI applications in drug discovery, focusing on the small data problem.
- Analyzing the limitations of traditional ML and DL models when applied to small datasets.
- Identifying and categorizing ML and DL methods tailored to overcome data scarcity in various DDD tasks.
Main Results:
- Limited data is a pervasive issue across multiple stages of the drug discovery and development pipeline.
- Traditional ML methods are suitable for small data but lack representational capacity, while DL methods risk overfitting.
- Adapted DL strategies and extended traditional ML approaches show promise for handling small datasets in DDD.
Conclusions:
- Addressing the small data problem is crucial for advancing AI in drug discovery and development.
- Integrating task-specific applications with methodological innovations is key to developing robust and generalizable AI models.
- Future research should focus on creating interpretable and trustworthy AI solutions that effectively leverage limited data in DDD.
More Related Videos
06:19Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018