Related Experiment Video
Updated: Nov 21, 2025

Author Spotlight: Enhancing Understanding and Treatment Strategies with the NEC-on-a-Chip Model
Published on: July 28, 2023
An ensemble learning approach for modeling the systems biology of drug-induced injury
Joaquim Aguirre-Plans1, Janet Piñero1, Terezinha Souza2
1Research Programme on Biomedical Informatics (GRIB), Hospital del Mar Medical Research Institute (IMIM), DCEXS, Pompeu Fabra University (UPF), Barcelona, Spain.
Background:
Drug-induced liver injury (DILI) is an adverse reaction caused by the intake of drugs of common use that produces liver damage. The impact of DILI is estimated to affect around 20 in 100,000 inhabitants worldwide each year. Despite being one of the main causes of liver failure, the pathophysiology and mechanisms of DILI are poorly understood. In the present study, we developed an ensemble learning approach based on different features (CMap gene expression, chemical structures, drug targets) to predict drugs that might cause DILI and gain a better understanding of the mechanisms linked to the adverse reaction.
Results:
We searched for gene signatures in CMap gene expression data by using two approaches: phenotype-gene associations data from DisGeNET, and a non-parametric test comparing gene expression of DILI-Concern and No-DILI-Concern drugs (as per DILIrank definitions). The average accuracy of the classifiers in both approaches was 69%. We used chemical structures as features, obtaining an accuracy of 65%. The combination of both types of features produced an accuracy around 63%, but improved the independent hold-out test up to 67%. The use of drug-target associations as feature obtained the best accuracy (70%) in the independent hold-out test.
Conclusions:
When using CMap gene expression data, searching for a specific gene signature among the landmark genes improves the quality of the classifiers, but it is still limited by the intrinsic noise of the dataset. When using chemical structures as a feature, the structural diversity of the known DILI-causing drugs hampers the prediction, which is a similar problem as for the use of gene expression information. The combination of both features did not improve the quality of the classifiers but increased the robustness as shown on independent hold-out tests. The use of drug-target associations as feature improved the prediction, specially the specificity, and the results were comparable to previous research studies.
Insights
Predicting drug-induced liver injury (DILI) is crucial. Drug-target associations offer the most accurate method for identifying DILI-causing drugs, improving understanding of liver damage mechanisms.
Area of Science:
- Pharmacology and Toxicology
- Computational Biology
- Drug Safety
Background:
- Drug-induced liver injury (DILI) is a significant adverse drug reaction causing liver damage, affecting approximately 20 in 100,000 people globally each year.
- Despite its prevalence and role in liver failure, the underlying pathophysiology and mechanisms of DILI remain poorly understood.
- Accurate prediction of DILI-causing drugs is essential for improving patient safety and understanding drug toxicity.
Purpose of the Study:
- To develop an ensemble learning approach for predicting drugs that may cause DILI.
- To investigate the utility of various features, including gene expression, chemical structures, and drug targets, in DILI prediction.
- To enhance the understanding of mechanisms associated with DILI.
Main Methods:
- Utilized Connectivity Map (CMap) gene expression data to identify gene signatures associated with DILI.
- Employed two approaches for gene signature identification: phenotype-gene associations and a non-parametric test comparing DILI-concern and no-DILI-concern drugs.
- Incorporated chemical structures and drug-target associations as features in ensemble learning models.
Main Results:
- Classifiers based on CMap gene expression achieved an average accuracy of 69%.
- Prediction models using chemical structures as features yielded an accuracy of 65%.
- The most accurate predictions were achieved using drug-target associations, with a 70% accuracy in independent hold-out tests.
Conclusions:
- Drug-target associations provided the best predictive performance, particularly in terms of specificity, comparable to existing research.
- Combining gene expression and chemical structure features improved model robustness but not overall accuracy.
- Further research into gene signatures and chemical structures is limited by data noise and structural diversity, respectively.
More Related Videos
Related Concept Videos
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Model Approaches for Pharmacokinetic Data: Physiological Models
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Mechanistic Models: Overview of Compartment Models
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...

