一种通用蛋白质识别方法,用于新和多样化的测序技术.
Bikash Kumar Bhandari1, Nick Goldman1
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB10 1SD, UK.
NAR genomics and bioinformatics
|September 19, 2024
概括
新的隐藏马尔科夫模型方法可以从杂的测序数据中准确识别蛋白质. 这种方法即使在具有有限氨基酸歧视的早期设备上也起作用,从而改善了蛋白质的发现.
科学领域:
- 生物化学 生物化学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 蛋白质测序技术正在迅速发展,但早期的设备会产生噪音和易出错的数据.
- 目前的方法难以从不完整或不准确的序列签名中识别蛋白质.
研究的目的:
- 开发一种广泛适用的蛋白质识别方法,使用来自下一代测序器的噪音序列数据.
- 为了评估这种方法在各种模拟的测序错误条件中的性能.
主要方法:
- 开发了一个隐藏的马尔科夫模型 (HMM) 来分析具有潜在错误的蛋白质签名.
- 该HMM方法在人类蛋白质数据库 (N=20,181) 上进行了测试,使用模拟新技术的假设测序装置进行了测试.
主要成果:
- 在各种条件下,HMM方法在识别蛋白质方面表现良好,包括不同的信号可分辨率和错误率 (插入,删除).
- 即使使用有限的氨基酸区分和碎片化的序列数据,也可以实现高精度.
结论:
- 开发的隐藏马尔科夫模型方法能够从杂的序列数据中准确识别蛋白质,即使使用早期测序设备.
- 这种方法预计将对未来广泛的蛋白质测序技术和应用有价值.
更多相关视频
相关概念视频
Peptide Identification Using Tandem Mass Spectrometry
6.4K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.4K
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Next-generation Sequencing
88.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.4K
Sanger Sequencing
753.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
753.8K


