Related Experiment Video
Updated: Oct 9, 2025

Fully Autonomous Characterization and Data Collection from Crystals of Biological Macromolecules
Published on: March 22, 2019
Random sampling of the Protein Data Bank: RaSPDB
1Department of Chemistry, University of Pavia, Viale Taramelli 12, 27100, Pavia, Italy. Oliviero.carugo@univie.ac.at.
Abstract:
A novel and simple procedure (RaSPDB) for Protein Data Bank mining is described. 10 PDB subsets, each containing 7000 randomly selected protein chains, are built and used to make 10 estimations of the average value of a generic feature F-the length of the protein chain, the amino acid composition, the crystallographic resolution, and the secondary structure composition. These 10 estimations are then used to compute an average estimation of F together with its standard error. It is heuristically verified that the dimension of these 10 subsets-7000 protein chains-is sufficiently small to avoid redundancy within each subset and sufficiently large to guarantee stable estimations amongst different subsets. RaSPDB has two major advantages over classical procedures aimed to build a single, non-redundant PDB subset: a larger fraction of the information stored in the PDB is used and an estimation of the standard error of F is possible.
More Related Videos
07:19Structural Studies of Macromolecules in Solution using Small Angle X-Ray Scattering
Published on: November 5, 2018
07:38Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019