Leveraging multi-source data to resolve inconsistency across pharmacogenomic datasets in drug sensitivity prediction

Xiaodi Li1, Trisha Das1,2, Kritib Bhattarai1,3

  • 1Department of Artificial Intelligence and Informatics Research, Mayo Clinic, Rochester, MN, USA.

Insights

This study introduces Aggregated Learning (AL), a computational model that improves drug response prediction by learning from inconsistencies across multiple pharmacogenomics datasets. The novel approach enhances model accuracy and generalizability for biomarker identification.

Area of Science:

  • Computational biology
  • Pharmacogenomics
  • Biomedical data science

Background:

  • Pharmacogenomics datasets are crucial for biomarker identification and drug response prediction.
  • Existing models often underperform due to inconsistencies stemming from inter-tumoral heterogeneity, experimental variations, and cell subtype complexity.
  • These inconsistencies limit the generalizability of predictive models.

Purpose of the Study:

  • To develop a computational model that enhances drug response prediction by effectively learning from inconsistencies across multiple pharmacogenomics datasets.
  • To improve the accuracy and generalizability of predictive models by addressing data variations.

Main Methods:

  • Proposed a novel computational model based on Aggregated Learning (AL).
  • The AL model was trained on overlapping inconsistent data points from three pharmacogenomic datasets: Cancer Cell Line Encyclopedia (CCLE), Genomics of Drug Sensitivity in Cancer (GDSC2), and cancer-genomics.org (gCSI).
  • Compared the AL model's performance against four baseline methods: Selecting Better (SB), Result Average (RA), Combining Data (CD), and Model Average (MA).

Main Results:

  • The Aggregated Learning (AL) model demonstrated superior performance compared to baseline methods.
  • Achieved lower Mean Absolute Error (MAE) scores: 0.090 for CCLE-GDSC, 0.096 for CCLE-gCSI, and 0.081 for GDSC-gCSI.
  • The results indicate that explicitly addressing dataset inconsistencies significantly enhances prediction accuracy.

Conclusions:

  • Addressing inconsistencies within and across pharmacogenomics datasets is critical for improving drug response prediction models.
  • The proposed Aggregated Learning (AL) model offers a promising solution for robust and generalizable drug response predictions.
  • This approach has the potential to advance personalized medicine through more accurate biomarker identification and treatment selection.

Related Concept Videos

Pharmacogenomics: Identification of New Drug Targets01:29

Pharmacogenomics: Identification of New Drug Targets

Advances in genomics have profoundly influenced drug discovery by increasing both the speed and accuracy of pharmaceutical development. Pharmacogenomics, which examines how genetic variation influences drug response, facilitates the identification of novel therapeutic targets and enables patient stratification for personalized treatment. These strategies contribute to improved drug efficacy, minimized adverse effects, and more efficient clinical trial design.Mapping genetic differences...
29
Pharmacogenetics and Pharmacogenomics: Overview01:29

Pharmacogenetics and Pharmacogenomics: Overview

Pharmacogenetics and pharmacogenomics examine how genetic factors influence an individual's response to drugs. While pharmacogenetics focuses on the impact of specific genetic variants on drug effects, pharmacogenomics takes a broader approach, studying how genetic variation across populations contributes to differences in drug responses. These fields aim to explain why individuals may experience varying levels of efficacy or adverse reactions to the same medication.Variability in drug...
51
Pharmacogenetics of Drug Metabolism: Overview01:27

Pharmacogenetics of Drug Metabolism: Overview

Genetic polymorphism in drug metabolism is crucial to the inter-individual variability observed in drug responses. Drug metabolism primarily involves the chemical modification of drugs and other xenobiotics to enhance their elimination by increasing their polarity. Two main classes of enzymes mediate this biotransformation process: Phase I enzymes, primarily cytochrome P450s, catalyze oxidation and reduction reactions, while other enzymes, such as esterases, mediate hydrolysis, and Phase II...
32
Analysis of Population Pharmacokinetic Data01:12

Analysis of Population Pharmacokinetic Data

Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
843
Principles of Pharmacogenetics: Types of Genetic Variants01:27

Principles of Pharmacogenetics: Types of Genetic Variants

The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
29
Pharmacogenetic Phenotypes: Alterations in Pharmacokinetics, Drug Targets and Biologic Milieu01:29

Pharmacogenetic Phenotypes: Alterations in Pharmacokinetics, Drug Targets and Biologic Milieu

Genetic variations significantly influence drug response through pharmacokinetics, receptor interactions, and biologic milieu modifications. Pharmacokinetic alterations impact drug metabolism and clearance, affecting efficacy and toxicity. Variants in drug-metabolizing enzymes, such as CYP2C9 and CYP2C19, alter drug activation and elimination. For example, CYP2C9 loss-of-function variants require lower warfarin doses to prevent excessive bleeding, while CYP2C19 variants reduce clopidogrel...
26