Related Experiment Videos
Machine learning models for adverse drug reaction prediction. A systematic review and three-level multilevel
Giorgio Cappello1, Christoph Oster1,2, Teresa Schmidt1
1Department of Neurology and Center for Translational Neuro- and Behavioral Sciences (C-TNBS), Division of Clinical Neurooncology, University Medicine Essen, University Duisburg-Essen, Essen, Germany.
Background:
Adverse drug reactions (ADRs) are a leading cause of preventable hospitalization, yet the dominant pharmacovigilance paradigm remains reactive. Machine learning (ML) offers a data-driven alternative, but the evidence base has not been formally meta-analyzed under clinically realistic inclusion criteria and a dependence-respecting statistical framework. This is the first PRISMA 2020-compliant and PROBAST-screened multilevel meta-analysis of broad-spectrum ML-based ADR prediction. We aimed to quantify pooled discrimination of ML models for broad-spectrum ADR prediction and to test the influence of algorithm class, protein-target features, and outcome breadth.
Methods:
Following a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 compliant, PROSPERO-registered protocol (CRD420250653686), we searched PubMed and Web of Science Core Collection from inception to 30 January 2025. A Scopus and IEEE Xplore top-up search on 2 February 2026 yielded no further records. Eligible studies applied ML to clinically validated data sources for multi-drug ADR prediction. Areas under the receiver operating characteristic curve (AUCs) were logit-transformed and pooled with a study-level DerSimonian-Laird random-effects model and a pre-specified primary three-level multilevel meta-regression (restricted maximum likelihood [REML]) using 30 model-level AUCs nested within 9 contributing studies.
Results:
Eleven Prediction model Risk Of Bias ASsessment Tool (PROBAST) screened studies were included, of which 9 contributed quantitatively. The pre-specified primary three-level multilevel pooled AUC was 0.841 (95% confidence interval [CI] 0.789-0.883), with a study-level DerSimonian-Laird sensitivity estimate of 0.832 (95% CI 0.788-0.869, I-squared 90.7%). Leave-one-study-out estimates ranged 0.818-0.842. Neural networks did not statistically outperform logistic regression, random forests, or k-nearest neighbor. An apparent protein-target feature penalty (model-level p = 0.0002) was driven by one dominant study and disappeared in multilevel sensitivity analysis (adjusted p = 0.75).
Conclusion:
ML models showed moderate discrimination under Grading of Recommendations Assessment, Development and Evaluation (GRADE) low certainty. Within this small and methodologically heterogeneous evidence base, the pooled AUC is a descriptive summary and does not establish broadly generalizable performance across data sources, ADR definitions, feature-engineering strategies, or validation settings. The high heterogeneity indicates substantial variation in underlying performance across contexts. Standardized benchmarks, harmonized outcome taxonomies, mandatory external validation, and regulator-aligned prospective evaluation are prerequisites for clinical deployment.
Systematic Review Registration:
[https://www.crd.york.ac.uk/PROSPERO/view/CRD420250653686], identifier [CRD420250653686].
Related Concept Videos
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Drug Toxicity: Risk factors
Pharmacodynamic Models: Additive and Proportional Drug Effect Model
Pharmacodynamic Models: Direct Effect Model and Indirect Response Model
Pharmaceutical Poisoning: Potential Scenarios