Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Peptide Identification Using Tandem Mass Spectrometry01:33

Peptide Identification Using Tandem Mass Spectrometry

8.8K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.8K
Tandem Mass Spectrometry01:21

Tandem Mass Spectrometry

2.8K
Tandem mass spectrometry is a technique that uses multiple mass analyzers in series to obtain a higher selectivity and reduce chemical noise during analyte detection. Instruments with multiple analyzers separated by an interaction cell enable secondary fragmentation and selected study of the fragment ions.Secondary fragmentations occur in the interaction cell and can be induced by various factors. Fragmentation induced by collision with inert gases, such as N2, Ar, He, etc., is called...
2.8K
Mass Spectrometry: Overview01:19

Mass Spectrometry: Overview

9.8K
Mass spectrometry is an analytical technique used to determine the molecular mass and molecular formula of a compound. The basic principle of mass spectrometry is to generate ions from the analyte molecule and measure these ion abundances against their molecular mass. One common type of ionization, known as electron ionization or EI, bombards the analyte molecules in the gas phase with high-energy electron beams. The electron beams displace an electron from the molecule and leave behind a...
9.8K
Mass Spectrum: Interpretation01:24

Mass Spectrum: Interpretation

3.8K
An unknown compound can be established by identifying the molecular ion peak in the mass spectrum. The molecular ion peak is often weak or absent due to the predominance of fragmentation in high-energy electron beams. In such cases, a soft-energy electron beam can be used to scan the spectrum to enhance the intensity of the molecular ion peak. Additionally, chemical ionization, field ionization, and desorption ionization spectra are used to obtain a relatively intense molecular ion peak.To...
3.8K
Mass Spectrometry: Complex Analysis01:21

Mass Spectrometry: Complex Analysis

2.0K
Mass spectrometry is an important technique for the identification of pure compounds. However, it has some limitations for the analysis of complex mixtures, often due to excessive fragmentation making the spectrum too complicated to decipher. Mass spectrometry can be combined with suitable separation methods in sequence, forming hyphenated methods, which are useful in the analysis of complex mixtures.
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
2.0K
Mass Spectrometry: Molecular Fragmentation Overview01:20

Mass Spectrometry: Molecular Fragmentation Overview

6.1K
The ionization of a molecule into a molecular ion inside the mass spectrometer causes instability in the molecule's structure due to the loss of an electron. This eventually leads to the fragmentation or breaking of some bonds in the molecule. The fragmentation occurs predominantly at specific bonds to yield relatively stable fragments.
One type of fragmentation pattern is the cleavage of a single bond in the molecular ion. The cleavage leads to a radical and a cation. The cleavage can occur at...
6.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Incorporating Scientific Knowledge into Neural Network Density Functionals.

Journal of chemical theory and computation·2026
Same author

KRASAVA-An Expert System for Virtual Screening of KRAS G12D Inhibitors.

International journal of molecular sciences·2026
Same author

Harmonic Scale Factors of Fundamental Transitions for Dispersion-corrected Quantum Chemical Methods.

Chemphyschem : a European journal of chemical physics and physical chemistry·2024
See all related articles

Related Experiment Video

Updated: Mar 15, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

70.0K

De Novo Structure Prediction from Tandem Mass Spectra: Algorithms, Benchmarks, and Limitations.

Mark Yu Schneider1, Daniil D Kholmanskikh1, Kirill Ya Romanov1

  • 1Research Center of the Artificial Intelligence Institute, Innopolis University, 420500 Innopolis, Russia.

Molecules (Basel, Switzerland)
|March 14, 2026
PubMed
Summary

Accurately identifying unknown molecules using de novo generative models is crucial for chemistry. Current models show low accuracy due to data leakage, highlighting the need for better benchmarking and formula-conditioned generation.

Keywords:
cheminformaticsde novo structure elucidationdiffusion modelsgenerative modelsmachine learningmass spectrometrymetabolomicstandem mass spectrometry

More Related Videos

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
10:37

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification

Published on: November 15, 2017

12.8K
Quaternary Structure Modeling Through Chemical Cross-Linking Mass Spectrometry: Extending TX-MS Jupyter Reports
05:18

Quaternary Structure Modeling Through Chemical Cross-Linking Mass Spectrometry: Extending TX-MS Jupyter Reports

Published on: October 20, 2021

2.7K

Related Experiment Videos

Last Updated: Mar 15, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

70.0K
Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
10:37

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification

Published on: November 15, 2017

12.8K
Quaternary Structure Modeling Through Chemical Cross-Linking Mass Spectrometry: Extending TX-MS Jupyter Reports
05:18

Quaternary Structure Modeling Through Chemical Cross-Linking Mass Spectrometry: Extending TX-MS Jupyter Reports

Published on: October 20, 2021

2.7K

Area of Science:

  • Analytical Chemistry
  • Computational Chemistry
  • Cheminformatics

Background:

  • Molecule identification from analytical data is vital for drug discovery and metabolomics.
  • Tandem mass spectrometry generates structural fingerprints, but most spectra lack library references.
  • De novo generative models aim to predict molecular structures but face accuracy assessment challenges.

Purpose of the Study:

  • To critically analyze the accuracy of state-of-the-art de novo generative models for molecule identification.
  • To investigate the impact of data leakage on reported model performance.
  • To review the evolution of generative model architectures and propose a roadmap for future development.

Main Methods:

  • Analysis of de novo generative models using leakage-controlled benchmarks (MassSpecGym).
  • Review of three architectural eras: RNNs, sequence models, and graph-native diffusion.
  • Evaluation of formula-conditioned vs. unconstrained generation and scaffold-based approaches.

Main Results:

  • State-of-the-art models exhibit only 4.1% top-10 accuracy on rigorously controlled benchmarks, contrary to prior reports.
  • Data leakage in naive data splits is identified as a primary cause for inflated performance metrics.
  • Explicitly conditioning models on molecular formulas significantly improves exact-match accuracy.

Conclusions:

  • Current de novo generative models for molecule identification require significant improvement in accuracy.
  • Standardized, leakage-aware benchmarking and transparent reporting are essential for reliable tool development.
  • Formula-conditioned generation shows promise for unknown discovery, while scaffold-based methods face bottlenecks with predicted scaffolds.