Related Experiment Video
Updated: May 20, 2026

Analyzing Multifactorial RNA-Seq Experiments with DiCoExpress
Published on: July 29, 2022
Predictive cheminformatics in drug discovery: statistical modeling for analysis of micro-array and gene expression
N Sukumar1, Michael P Krein, Mark J Embrechts
1Rensselaer Exploratory Center for Cheminformatics Research and Department of Chemistry and Chemical Biology, Rensselaer Polytechnic Institute, Troy, NY, USA. nagams@rpi.edu
Predictive cheminformatics and bioinformatics use statistical modeling to analyze vast chemical and biological data for drug discovery. Best practices ensure accurate data analysis, model validation, and interpretation for reliable results.
Area of Science:
- Computational chemistry and biology
- Data science and machine learning
- Drug discovery and development
Background:
- High-throughput screening and microarray technologies generate massive chemical and biological datasets.
- Computational techniques are essential for analyzing, visualizing, and modeling this complex data.
- Predictive cheminformatics and bioinformatics leverage statistical methods to extract valuable insights from large-scale biological and chemical data.
Purpose of the Study:
- To review critical considerations for applying statistical methods in predictive cheminformatics and bioinformatics.
- To summarize best practices for data representation, preprocessing, and model validation.
- To guide the effective use of statistical modeling in drug development through data mining.
Main Methods:
- Review of current statistical modeling techniques in cheminformatics and bioinformatics.
- Discussion of data representation, preprocessing, and model validation strategies.
- Analysis of domain applicability, similarity assessment, and model interpretation in predictive modeling.
Main Results:
- Identification of key factors influencing the success of predictive modeling in drug discovery.
- Emphasis on the importance of rigorous validation and interpretation of statistical models.
- Highlighting the need for careful consideration of data characteristics and structure-activity relationships.
Conclusions:
- Effective application of statistical methods requires attention to data quality, preprocessing, and validation.
- Best practices in predictive cheminformatics enhance the reliability of insights derived from large datasets.
- Adherence to these principles is crucial for successful drug development using computational approaches.
Related Concept Videos
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
DNA Microarrays
Statistical Software for Data Analysis and Clinical Trials
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This relationship...
Pharmacogenomics: Identification of New Drug Targets