Related Experiment Video
Updated: Feb 7, 2026

Patient-Derived Tumor Explants As a "Live" Preclinical Platform for Predicting Drug Resistance in Patients
Published on: February 7, 2021
Learning predictive models of drug side-effect relationships from distributed representations of literature-derived
Justin Mower1, Devika Subramanian2, Trevor Cohen3
1Baylor College of Medicine, Quantitative and Computational Biosciences, Houston, Texas, USA.
Objective:
The aim of this work is to leverage relational information extracted from biomedical literature using a novel synthesis of unsupervised pretraining, representational composition, and supervised machine learning for drug safety monitoring.
Methods:
Using ≈80 million concept-relationship-concept triples extracted from the literature using the SemRep Natural Language Processing system, distributed vector representations (embeddings) were generated for concepts as functions of their relationships utilizing two unsupervised representational approaches. Embeddings for drugs and side effects of interest from two widely used reference standards were then composed to generate embeddings of drug/side-effect pairs, which were used as input for supervised machine learning. This methodology was developed and evaluated using cross-validation strategies and compared to contemporary approaches. To qualitatively assess generalization, models trained on the Observational Medical Outcomes Partnership (OMOP) drug/side-effect reference set were evaluated against a list of ≈1100 drugs from an online database.
Results:
The employed method improved performance over previous approaches. Cross-validation results advance the state of the art (AUC 0.96; F1 0.90 and AUC 0.95; F1 0.84 across the two sets), outperforming methods utilizing literature and/or spontaneous reporting system data. Examination of predictions for unseen drug/side-effect pairs indicates the ability of these methods to generalize, with over tenfold label support enrichment in the top 100 predictions versus the bottom 100 predictions.
Discussion And Conclusion:
Our methods can assist the pharmacovigilance process using information from the biomedical literature. Unsupervised pretraining generates a rich relationship-based representational foundation for machine learning techniques to classify drugs in the context of a putative side effect, given known examples.
More Related Videos
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Related Concept Videos
Drug Distribution: Volume of Distribution
Drug Distribution: Overview
Initially, the free drug in the...
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Drug Distribution: Tissue Binding
For...
Drug Distribution as One-Compartment Model and Elimination by Nonlinear Pharmacokinetics: Overview
For instance, consider the metabolism of sodium salicylate. This compound is metabolized into two distinct substances: a glucuronide and a glycine conjugate. The rate of conjugation depends...
State Space Representation
Consider an RLC circuit, a...