Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Protein Families02:47

Protein Families

16.0K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
16.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

What Is the Crystallographic Resolution of Structural Models of Proteins Generated with AlphaFold2?

ACS chemical biology·2024
Same author

Location of S-nitrosylated cysteines in protein three-dimensional structures.

Proteins·2023
Same author

Chalcogen bonds formed by protein sulfur atoms in proteins. A survey of high-resolution structures deposited in the protein data bank.

Journal of biomolecular structure & dynamics·2022
Same author

Interplay between hydrogen and chalcogen bonds in cysteine.

Proteins·2022
Same author

Survey of the Intermolecular Disulfide Bonds Observed in Protein Crystal Structures Deposited in the Protein Data Bank.

Life (Basel, Switzerland)·2022
Same author

B-factor accuracy in protein crystal structures.

Acta crystallographica. Section D, Structural biology·2022

Related Experiment Video

Updated: Oct 9, 2025

Fully Autonomous Characterization and Data Collection from Crystals of Biological Macromolecules
07:11

Fully Autonomous Characterization and Data Collection from Crystals of Biological Macromolecules

Published on: March 22, 2019

7.0K

Random sampling of the Protein Data Bank: RaSPDB.

Oliviero Carugo1,2

  • 1Department of Chemistry, University of Pavia, Viale Taramelli 12, 27100, Pavia, Italy. Oliviero.carugo@univie.ac.at.

Scientific Reports
|December 18, 2021
PubMed
Summary

A new method, RaSPDB, analyzes Protein Data Bank (PDB) data using multiple subsets. This approach provides more accurate estimations of protein features and their variability.

More Related Videos

Structural Studies of Macromolecules in Solution using Small Angle X-Ray Scattering
07:19

Structural Studies of Macromolecules in Solution using Small Angle X-Ray Scattering

Published on: November 5, 2018

12.9K
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.9K

Related Experiment Videos

Last Updated: Oct 9, 2025

Fully Autonomous Characterization and Data Collection from Crystals of Biological Macromolecules
07:11

Fully Autonomous Characterization and Data Collection from Crystals of Biological Macromolecules

Published on: March 22, 2019

7.0K
Structural Studies of Macromolecules in Solution using Small Angle X-Ray Scattering
07:19

Structural Studies of Macromolecules in Solution using Small Angle X-Ray Scattering

Published on: November 5, 2018

12.9K
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.9K

Area of Science:

  • Structural biology
  • Bioinformatics
  • Computational biology

Background:

  • The Protein Data Bank (PDB) is a crucial resource for structural biology.
  • Analyzing large PDB datasets for specific features can be computationally intensive and prone to redundancy.
  • Classical PDB subsetting methods may not fully utilize available information or provide error estimations.

Purpose of the Study:

  • To introduce a novel and simple procedure, RaSPDB, for mining the Protein Data Bank.
  • To develop a method for estimating average protein features and their standard errors from PDB data.
  • To compare the efficiency and information utilization of RaSPDB against classical PDB subsetting approaches.

Main Methods:

  • Generation of 10 PDB subsets, each with 7000 randomly selected protein chains.
  • Estimation of a generic feature (F) across these 10 subsets, including protein chain length, amino acid composition, crystallographic resolution, and secondary structure composition.
  • Computation of an average estimation for feature F and its standard error using the 10 subset estimations.

Main Results:

  • The RaSPDB method utilizes a subset size of 7000 protein chains, balancing redundancy avoidance and stable estimation.
  • The procedure allows for the estimation of the standard error for the analyzed protein features.
  • RaSPDB effectively uses a larger fraction of information stored in the PDB compared to single-subset methods.

Conclusions:

  • RaSPDB offers a robust and efficient approach for PDB data mining.
  • The method provides valuable insights into the variability and statistical properties of protein features.
  • RaSPDB enhances the analysis of structural biology data by enabling more comprehensive information extraction and error quantification.