PEPMatch:一种工具,用于识别大量蛋白质中的短序列匹配
Daniel Marrama1, William D Chronister1, Luise Westernberg1
1Division of Vaccine Discovery, La Jolla Institute for Immunology, La Jolla, San Diego, CA, USA.
BMC bioinformatics
|December 19, 2023
概括
PEPMatch显著加快了在大型蛋白质数据库中搜索短序列的速度,例如T细胞表位. 这种新工具在不牺牲精度的情况下,比BLAST提高了50倍的速度.
科学领域:
- 生物信息学是一种生物信息学.
- 免疫学 免疫学 免疫学
- 计算生物学 计算生物学
背景情况:
- 现有的生物序列比较工具对于短匹配缺乏效率.
- 识别线性T细胞表位 (8-15残留物) 对免疫学应用至关重要.
- 应用包括病毒表位保护,过敏原识别和癌症新表位分析.
研究的目的:
- 开发和评估一种用于快速准确的短序列匹配的工具.
- 为评估匹配工具建立基准.
- 解决在大型蛋白质基因数据集中有效地进行表位标识的需求.
主要方法:
- 为实用的短序相匹配场景定义了基准.
- 评估了现有的速度和回忆的方法.
- 开发了PEPMatch,这是一个使用确定性k-mer映射算法与蛋白质组预处理的新工具.
主要成果:
- 与BLAST相比,PEPMatch实现了50倍的速度增加.
- 该工具保持了高回忆率,确保了序列匹配的准确性.
- 基准数据集和PEPMatch代码是公开的.
结论:
- PEPMatch为酸序列匹配提供了相当大的速度和回忆优势.
- 开发的基准测试框架为未来的工具评估提供了一个标准.
- 免疫学家和研究人员可以轻松访问PEPMatch.
相关概念视频
Peptide Identification Using Tandem Mass Spectrometry
6.5K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.5K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K


