Papyrus: a large-scale curated dataset aimed at bioactivity predictions

O J M Béquignon1, B J Bongers1, W Jespers1

  • 1Division of Drug Discovery and Safety, Leiden Academic Centre for Drug Research, Leiden University, Leiden, The Netherlands.

Summary

The Papyrus dataset offers 60 million ligand-protein bioactivity data points, standardized for machine learning. This valuable resource aims to streamline predictive modeling for researchers by providing an accessible, high-quality benchmark dataset.