Related Experiment Video
Updated: Aug 15, 2025

Network Pharmacology Prediction and Experimental Validation of Trichosanthes-Fritillaria thunbergii Action Mechanism Against Lung Adenocarcinoma
Published on: March 3, 2023
Papyrus: a large-scale curated dataset aimed at bioactivity predictions
O J M Béquignon1, B J Bongers1, W Jespers1
1Division of Drug Discovery and Safety, Leiden Academic Centre for Drug Research, Leiden University, Leiden, The Netherlands.
The Papyrus dataset offers 60 million ligand-protein bioactivity data points, standardized for machine learning. This valuable resource aims to streamline predictive modeling for researchers by providing an accessible, high-quality benchmark dataset.
Area of Science:
- Computational chemistry
- Bioinformatics
- Machine learning in drug discovery
Background:
- Publicly available ligand-protein bioactivity data is rapidly growing, presenting opportunities for machine learning.
- Data quality and accessibility challenges hinder the effective use of this data for research.
- Researchers spend significant time adapting and identifying suitable datasets for their specific needs.
Purpose of the Study:
- To construct a comprehensive and standardized dataset for machine learning in drug discovery.
- To address the challenges of data heterogeneity, quality, and accessibility in bioactivity data.
- To provide a benchmark dataset for developing predictive models and facilitating research.
Main Methods:
- Aggregation of multiple large public datasets (e.g., ChEMBL, ExCAPE-DB) with smaller, high-quality datasets.
- Standardization and normalization of approximately 60 million data points for machine learning suitability.
- Demonstration of data filtering techniques and application in quantitative structure-activity relationship (QSAR) and proteochemometric modeling.
Main Results:
- Creation of the Papyrus dataset, comprising ~60 million standardized ligand-protein bioactivity data points.
- Successful integration and harmonization of diverse public bioactivity datasets.
- Illustrative examples of data filtering, QSAR, and proteochemometric modeling using the Papyrus dataset.
Conclusions:
- The Papyrus dataset serves as a valuable, accessible benchmark for constructing predictive models in drug discovery.
- Standardized and curated bioactivity data significantly enhances the efficiency of machine learning applications.
- This resource aims to accelerate research by providing a reliable foundation for computational studies.
More Related Videos
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
08:15Author Spotlight: Network Pharmacology and Molecular Docking to Decipher the Action of Jiawei Shengjiang San Against Diabetic Kidney Disease
Published on: May 10, 2024
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Globular and Fibrous Proteins
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...