Related Experiment Video
Updated: May 13, 2025

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
TopCysteineDB: A Cysteinome-wide Database Integrating Structural and Chemoproteomics Data for Cysteine Ligandability
Michele Bonus1, Julian Greb1, Jaimeen D Majmudar2
1Institute for Pharmaceutical and Medicinal Chemistry, Heinrich Heine University Düsseldorf 40225 Düsseldorf, Germany.
Abstract:
Development of targeted covalent inhibitors and covalent ligand-first approaches have emerged as a powerful strategy in drug design, with cysteines being attractive targets due to their nucleophilicity and relative scarcity. While structural biology and chemoproteomics approaches have generated extensive data on cysteine ligandability, these complementary data types remain largely disconnected. Here, we present TopCysteineDB, a comprehensive resource integrating structural information from the PDB with chemoproteomics data from activity-based protein profiling experiments. Analysis of the complete PDB yielded 264,234 unique cysteines, while the proteomics dataset encompasses 41,898 detectable cysteines across the human proteome. Using TopCovPDB, an automated classification pipeline complemented by manual curation, we identified 787 covalent cysteines and systematically categorized other functional roles, including metal-binding, cofactor-binding, and disulfide bonds. Mapping residue-wise structural information to sequence space enabled cross-referencing between structural and proteomics data, creating a unified view of cysteine ligandability. For TopCySPAL, a machine learning model was developed, integrating structural features and proteomics data, achieving strong predictive performance (AUROC: 0.964, AUPRC: 0.914) and robust generalization to novel cases. TopCysteineDB and TopCySPAL are freely accessible through a webinterface, TopCysteineDBApp (https://topcysteinedb.hhu.de/), designed to facilitate exploration of cysteine sites across the human proteome. The interface provides an interactive visualization featuring a color-coded mapping of chemoproteomics data onto cysteine site structures and the highlighting of identified peptide sequences. It offers customizable dataset downloads and ligandability predictions for user-provided structures. This resource advances targeted covalent inhibitor design by providing integrated access to previously dispersed data types and enabling systematic analysis and prediction of cysteine ligandability.
More Related Videos
11:19Label-Free Immunoprecipitation Mass Spectrometry Workflow for Large-scale Nuclear Interactome Profiling
Published on: November 17, 2019
19:16The Importance of Correct Protein Concentration for Kinetics and Affinity Determination in Structure-function Analysis
Published on: March 17, 2010
Related Concept Videos
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Ligand Binding and Linkage
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...