Related Experiment Video
Updated: Mar 28, 2026

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Logistic regression models to predict solvent accessible residues using sequence- and homology-based qualitative and
Reecha Nepal1, Joanna Spencer2, Guneet Bhogal3
1Department of Chemistry, San Jose State University , San Jose, CA 95192-0101, USA.
This study presents logistic regression models for predicting protein relative solvent accessibility (RSA). The models achieve accurate predictions using sequence and disorder propensity descriptors, aiding researchers in exploring RSA prediction applications.
Area of Science:
- Computational Biology
- Structural Bioinformatics
- Machine Learning in Biology
Background:
- Predicting protein relative solvent accessibility (RSA) is crucial for understanding protein structure and function.
- Existing methods for RSA prediction vary in computational intensity and accuracy.
- Developing accurate and computationally efficient RSA prediction models remains an active research area.
Purpose of the Study:
- To develop and validate novel logistic regression models for predicting protein residue solvent accessibility.
- To evaluate the performance of models incorporating sequence-based and disorder propensity descriptors.
- To provide accessible tools for researchers to explore RSA prediction.
Main Methods:
- Utilized a domain-complete learning set of over 1300 proteins.
- Developed logistic regression models using qualitative (amino acid type) and quantitative (sequence entropy) descriptors.
- Incorporated homology descriptors from BLASTp alignments and residue disorder propensity.
- Trained models using dichotomous responses (buried/accessible) derived from RSA values from crystal structures.
Main Results:
- Achieved binary prediction accuracies comparable to existing methods using standard RSA thresholds (20% and 25%).
- Demonstrated incremental accuracy improvements by including the Lobanov-Galzitskaya residue disorder propensity descriptor.
- Reported 25% threshold accuracies of 76.12% and 74.79% on the Manesh-215 and CASP(8+9) test sets, respectively.
Conclusions:
- Logistic regression models with sequence and disorder propensity descriptors provide accurate protein solvent accessibility predictions.
- The developed models offer a computationally efficient and interpretable approach to RSA prediction.
- The provided software and datasets facilitate further research and educational exploration in RSA prediction.
More Related Videos
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Entropy and Solvation
Predicting Molecular Geometry
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Chemical Shift: Internal References and Solvent Effects
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...

