Related Experiment Video
Updated: Jun 1, 2026

Capillary Electrophoresis Mass Spectrometry Approaches for Characterization of the Protein and Metabolite Corona Acquired by Nanomaterials
Published on: October 27, 2020
Determinants of protein corona adsorption and abundance revealed by interpretable machine learning across
Keyuan Li1, Alexa Canchola1, Fan Zhang2
1Environmental Toxicology Graduate Program, College of Natural & Agricultural Sciences, University of California, Riverside, CA, 92521, USA.
Abstract:
Nanoparticles (NPs) hold significant potential in biotechnology, including molecular sensing, controlled release systems, and therapeutic applications. However, their behavior in biological environments remains difficult to predict because proteins rapidly absorb onto NP surfaces, forming a protein corona (PC) that reshapes their surface properties and determines their biological identity, transport, and cellular interactions. In this study, we developed large-scale deep neural network (DNN) models to predict both protein adsorption (binary classification) and relative protein abundance (regression) on NP surfaces. We utilized a well-curated and comprehensive PC dataset comprising data from 83 peer-reviewed studies, 817 NP-PC samples, and 2,497 proteins, substantially expanding the scale and diversity compared with prior studies. Then, we employed a prevalence-based filtering strategy to mitigate sparsity and batch noise and trained over 200 machine learning models across proteins. The adsorption classification models achieved high discriminative performance (AUC = 0.96), while the abundance models achieved a pooled R² of 0.67 and an average per-protein R2 of 0.40 on the test set. SHapley Additive exPlanations (SHAP) revealed that adsorption was predominantly governed by NP material class and surface chemistry, whereas abundance was more strongly influenced by experimental handling and kinetic parameters, particularly isolation and incubation time. Incorporation of applicability domain (AD) analysis enabled identification of reliable prediction regions, with in-AD predictions demonstrating higher confidence and reduced error for both tasks. Together, these results demonstrate that our DNN models can identify predictive drivers of PC composition and reveal feature associations consistent with patterns reported in prior mechanistic literature, offering a data-driven reference to inform nanomaterial design for biomedical and environmental applications.
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence.
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein-Drug Binding: Determination Methods
Indirect methods involve isolating the bound drug from its free form in biological samples such as blood, serum, or plasma. These techniques aim to measure the percentage of drugs bound to proteins. Equilibrium dialysis is a commonly used method where the free drug concentration at equilibrium is measured by separating the bound...
Factors Affecting Protein-Drug Binding: Protein-Related Factors
The physicochemical properties of a drug play a significant role in its ability to bind to proteins. Lipophilic drugs, which dissolve in fats, oils, and lipids, can be bound by...

