Improved de novo peptide sequencing using LC retention time information
Yves Frank1, Tomas Hruz1, Thomas Tschager1
1Department of Computer Science, ETH Zurich, Universitätstrasse 6, 8092 Zürich, Switzerland.
Algorithms for Molecular Biology : AMB
|September 6, 2018
Summary
This study enhances de novo peptide sequencing by integrating liquid chromatography retention time with mass spectrometry data. Exploiting chromatographic information significantly improves peptide identification rates in proteomics.
Area of Science:
- Proteomics
- Analytical Chemistry
- Bioinformatics
Background:
- Liquid chromatography-tandem mass spectrometry (LC-MS/MS) is crucial for peptide identification in proteomics.
- De novo peptide sequencing reconstructs amino acid sequences from MS/MS data.
- Existing algorithms primarily use mass spectrum data, neglecting chromatographic information.
Purpose of the Study:
- To develop de novo peptide sequencing algorithms that incorporate liquid chromatography retention time data.
- To improve the accuracy and identification rates of peptide sequencing.
Main Methods:
- Developed novel algorithms for de novo peptide sequencing.
- Integrated three distinct models for predicting peptide retention time.
- Algorithms were designed to consider both mass spectrum and retention time data.
Main Results:
- Incorporating liquid chromatography retention time improves peptide identification rates.
- Evaluated algorithms on experimental data from synthesized peptides.
- Demonstrated enhanced performance compared to methods using only mass spectra.
Conclusions:
- Exploiting chromatographic information is beneficial for de novo peptide sequencing.
- The proposed methods offer improved accuracy in reconstructing peptide sequences.
- This approach advances the capabilities of LC-MS/MS in proteomic analysis.
Related Concept Videos
Peptide Bonds
83.2K
A peptide bond covalently attaches amino acids through a dehydration reaction. One amino acid's carboxyl group and another amino acid's amino group combine, releasing a water molecule. The resulting bond is the peptide bond. The products that such linkages form are peptides. As more amino acids join this growing chain, the resulting chain is a polypeptide. Each polypeptide has a free amino group at one end. This end has the N-terminal, or the amino-terminal, and the other end has a free...
83.2K
Cis-regulatory Sequences
11.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.8K
Bioavailability Enhancement: Drug Stability Enhancement and GI Retention
222
Body:Improving a drug's stability in the gastrointestinal (GI) tract is paramount for enhancing its bioavailability and therapeutic effectiveness. Various strategies are employed to protect the drug from the harsh gastric milieu and to ensure its release and absorption at the desired site within the GI tract.Polymer coatings are one such method used to shield drugs from the stomach's acidic environment. By preventing premature drug release, these coatings improve the bioavailability of unstable...
222
Sequences
278
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
278
Sanger Sequencing
774.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.7K
Arithmetic Sequences
240
An arithmetic sequence is a structured arrangement of numbers where each term is derived by adding a constant value, known as the common difference, to the previous term. This consistent pattern allows for the efficient computation of any term within the sequence as well as the cumulative sum of multiple terms. The formula for finding the nth term of an arithmetic sequence is:Here, aₙ represents the nth term of the sequence, a is the first term, d is the common difference, and n is the...
240


