Related Experiment Video
Updated: Aug 2, 2026

10:53
Membrane-SPINE: A Biochemical Tool to Identify Protein-protein Interactions of Membrane Proteins In Vivo
Published on: November 8, 2013
SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput
1Department of Molecular, Cellular and Developmental Biology, Yale University, New Haven, CT 06520, USA.
Nucleic Acids Research
|July 4, 2001
Summary
The SPINE database standardizes structural proteomics data for analysis and collaboration. Machine learning predicts protein solubility and crystallization from sequence features, identifying key rules for improved experimental success.
Area of Science:
- Structural Proteomics
- Bioinformatics
- Computational Biology
- Machine Learning in Biology
Background:
- High-throughput structural proteomics generates vast amounts of protein structure determination data.
- Standardizing this data is crucial for retrospective analysis and data mining.
- Existing systems often lack robust features for systematic analysis and collaboration.
Purpose of the Study:
- To develop a database and analysis system (SPINE) for standardizing and mining structural proteomics data.
- To enable distributed scientific collaboration via the internet.
- To apply machine learning to predict protein properties like solubility and crystallization propensity.
Main Methods:
- Creation of the SPINE database and analysis system for the Northeast Structural Genomics Consortium.
- Development of specifications and ontologies for data standardization.
- Application of machine learning, specifically decision trees, using sequence-derived features to predict protein solubility and crystallization.
Main Results:
- The SPINE database currently holds experimental data on 985 constructs from various organisms.
- Machine learning models successfully predicted protein solubility and crystallization propensity.
- Key predictive rules identified include that soluble proteins tend to have more acidic residues and fewer hydrophobic stretches.
Conclusions:
- The SPINE system effectively standardizes proteomics data, enabling systematic data mining and collaborative research.
- Sequence-based machine learning models can accurately predict crucial experimental outcomes like protein solubility and crystallization.
- The study highlights the utility of decision trees and proposes novel error estimation methods for intermediate-sized datasets.

