Related Experiment Video
Updated: Sep 3, 2026

TMT Sample Preparation for Proteomics Facility Submission and Subsequent Data Analysis
Published on: June 8, 2020
How Reusable Are Public Proteomics Datasets in Practice?
Giuseppe G F Leite1, Victor Corasolla Carregari2, Alexandre Keiji Tashima3,4
1Division of Infectious Diseases, Department of Medicine, Escola Paulista de Medicina, Universidade Federal de Sao Paulo, São Paulo, Brazil.
Abstract:
Public proteomics repositories have expanded rapidly, but large-scale reuse still depends on whether deposited studies are computationally reusable. We provide an evidence-based snapshot of this gap by evaluating 500 non-redundant human ProteomeXchange datasets released in 2025. Experimental design annotation was classified as metadata-based when a deposited table (e.g., SDRF, spreadsheet, or text table) explicitly linked samples, files, channels, or quantitative columns to biological groups; sample-based when groups could only be inferred from sample, column or channel names; and no information when reliable group assignment was not possible. Sample-based datasets represented the largest category, comprising 195 deposits (39.0%), followed by datasets lacking reliable design information (170; 34.0%) and metadata-based datasets (135; 27.0%). Among metadata-based deposits, 81 (60.0%) included additional contextual covariates. Processed quantitative outputs were fully available in 245 datasets (49.0%), partially available in 168 (33.6%), and absent in 87 (17.4%). Combining processed quantitative outputs with available biological design information, 189 datasets (37.8%) supported direct matrix-level reuse. Under a stricter scenario requiring metadata-based design, contextual covariates and available processed quantitative outputs, only 64 datasets (12.8%) fulfilled all criteria. These findings suggest that scalable reuse requires more consistent machine-actionable sample-to-file mapping, interpretable design annotation, processing information and processed outputs, while preserving raw data for harmonized reanalysis.

