Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genomics02:02

Genomics

41.2K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
41.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Chain-Mobility-Enabled 1D Lanthanide Coordination Polymer Glassy Scintillators for Underwater X-Ray Videography.

Advanced materials (Deerfield Beach, Fla.)·2026
Same author

Beyond Heuristics: A Model-Agnostic Framework for Uncertainty Quantification in QSAR via Adaptive Conformal Prediction.

Chemical research in toxicology·2026
Same author

EC-Dock: A Fast Equivariant Consistency Model for Molecular Docking and Virtual Screening.

Journal of chemical information and modeling·2026
Same author

The changing epidemiology of human type 2 diabetes-associated atherosclerosis: Pathophysiological mechanisms and emerging treatment possibilities.

Journal of internal medicine·2026
Same author

Stimuli-Responsive Triplet Emission and X-Ray Scintillation via Reversible Structural Switching in Pyromellitic Diimide Cocrystals.

Angewandte Chemie (International ed. in English)·2026
Same author

Integrating Machine Learning Interatomic Potentials with MMPBSA for Accurate Protein-Ligand Binding Free Energy Calculations.

The journal of physical chemistry. B·2026

Related Experiment Video

Updated: Mar 6, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.7K

ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics.

Jiangming Sun1, Nina Jeliazkova2, Vladimir Chupakin3

  • 1Discovery Sciences, Innovative Medicines and Early Development Biotech Unit, AstraZeneca R&D Gothenburg, 43183 Mölndal, Sweden.

Journal of Cheminformatics
|March 21, 2017
PubMed
Summary

This study compiles a large chemogenomics dataset from public databases, containing over 70 million structure-activity relationship data points. This resource supports building accurate in silico target prediction and polypharmacology models.

Keywords:
Big DataBioactivityChemical structureChemogenomicsMolecular fingerprintsQSARSearch engine

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K
A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
07:40

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions

Published on: May 27, 2021

4.7K

Related Experiment Videos

Last Updated: Mar 6, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.7K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K
A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
07:40

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions

Published on: May 27, 2021

4.7K

Area of Science:

  • Computational chemistry and cheminformatics
  • Drug discovery and development
  • Bioinformatics and data science

Background:

  • Chemogenomics data is crucial for developing in silico target prediction models.
  • The growing volume of chemogenomics data presents opportunities for Big Data approaches.
  • High-quality datasets are essential for effective model building.

Purpose of the Study:

  • To compile a comprehensive, high-quality chemogenomics dataset.
  • To create an industry-scale resource for computational drug discovery.
  • To facilitate the development and validation of in silico prediction models.

Main Methods:

  • Integrated data from public databases (PubChem, ChEMBL).
  • Collected over 70 million structure-activity relationship (SAR) data points.
  • Included chemical structures, target information, and activity annotations.

Main Results:

  • A large-scale chemogenomics dataset with over 70 million SAR data points.
  • A valuable resource for in silico polypharmacology and off-target effect prediction.
  • A foundation for validating cheminformatics approaches.

Conclusions:

  • The compiled dataset is a significant resource for advancing computational drug discovery.
  • Enables the development of more robust predictive models.
  • Supports the validation of various cheminformatics methodologies.