Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Peptide Identification Using Tandem Mass Spectrometry01:33

Peptide Identification Using Tandem Mass Spectrometry

6.4K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.4K
Mass Spectrometry: Complex Analysis01:21

Mass Spectrometry: Complex Analysis

725
Mass spectrometry is an important technique for the identification of pure compounds. However, it has some limitations for the analysis of complex mixtures, often due to excessive fragmentation making the spectrum too complicated to decipher. Mass spectrometry can be combined with suitable separation methods in sequence, forming hyphenated methods, which are useful in the analysis of complex mixtures.
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
725

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A generalizable Hi-C foundation model for chromatin architecture, single-cell and multiomics analysis across species.

Nature methods·2026
Same author

Revisiting resonance-excitation collision-induced dissociation for data-independent acquisition.

bioRxiv : the preprint server for biology·2026
Same author

Label-Free Quantification in the Crux Toolkit.

Journal of proteome research·2026
Same author

Prioritizing peptides for targeted mass spectrometry experiments using deep learning.

bioRxiv : the preprint server for biology·2026
Same author

Embryo-scale Visual Cell Sorting reveals a conserved transcriptomic signature of nucleolar size linked to proteostasis.

bioRxiv : the preprint server for biology·2026
Same author

A quantitative proteomics dataset for assessment and prediction of low dose X-ray radiation exposure in mice.

bioRxiv : the preprint server for biology·2026

Related Experiment Video

Updated: Jun 8, 2025

Comprehensive Workflow of Mass Spectrometry-based Shotgun Proteomics of Tissue Samples
14:51

Comprehensive Workflow of Mass Spectrometry-based Shotgun Proteomics of Tissue Samples

Published on: November 13, 2021

5.2K

A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models.

Bo Wen1, William Stafford Noble2,3

  • 1Department of Genome Sciences, University of Washington, Seattle, WA, USA.

Scientific Data
|November 8, 2024
PubMed
Summary

A large dataset of 2.8 million high-confidence peptide-spectrum matches from nine species was created. This resource supports training machine learning models for proteomics tasks like de novo sequencing.

More Related Videos

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
10:37

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification

Published on: November 15, 2017

11.9K
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.7K

Related Experiment Videos

Last Updated: Jun 8, 2025

Comprehensive Workflow of Mass Spectrometry-based Shotgun Proteomics of Tissue Samples
14:51

Comprehensive Workflow of Mass Spectrometry-based Shotgun Proteomics of Tissue Samples

Published on: November 13, 2021

5.2K
Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
10:37

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification

Published on: November 15, 2017

11.9K
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.7K

Area of Science:

  • Proteomics
  • Bioinformatics
  • Machine Learning

Background:

  • Machine learning model training for de novo sequencing and spectral clustering requires extensive, high-confidence peptide-spectrum match datasets.
  • Existing benchmarks may lack consistent data quality or proper separation of training and testing data.

Purpose of the Study:

  • To introduce a large-scale, high-quality dataset of peptide-spectrum matches for machine learning applications in proteomics.
  • To provide a reliable resource for training and evaluating machine learning models in sequence analysis.

Main Methods:

  • Re-processing of a previously described benchmark dataset.
  • Ensuring consistent data quality across all spectra.
  • Strict separation of training and testing peptides to prevent data leakage.

Main Results:

  • A dataset comprising 2.8 million high-confidence peptide-spectrum matches.
  • Data derived from nine diverse species.
  • The dataset is optimized for machine learning model development and validation.

Conclusions:

  • The newly processed dataset offers a robust foundation for advancing machine learning in proteomics.
  • This resource facilitates the development of more accurate de novo sequencing and spectral clustering algorithms.