伪序列在新抗原预测中是否重要?
bioRxiv : the preprint server for biology
|December 22, 2025
概括
在个性化癌症疫苗中,生物知情的伪序列对于准确的新抗原预测至关重要. 精选的伪序列,特别是30-35残留的伪序列,仍然优于预测T细胞反应的替代编码.
科学领域:
- 免疫信息学是指免疫信息学.
- 计算生物学是一种计算生物学.
- 癌症免疫学 癌症免疫学
背景情况:
- 个性化癌症疫苗依赖于预测触发T细胞反应的新抗原.
- 目前的方法经常使用MHC I类伪序列,但它们的最佳定义和编码尚不清楚.
研究的目的:
- 系统地评估不同的主要基因相容性复合体 (MHC) I类表征,用于新抗原预测.
- 为了比较伪序列策略,包括基于结构,进化,随机和不同长度的伪序列策略,以及来自蛋白质语言模型和图形注释的嵌入.
主要方法:
- 利用BigMHC EL框架来评估各种MHC等位基因表征.
- 比较生物知情伪序列 (结构,进化多样性) 与随机基线和不同长度的性能.
- 从ESM-2蛋白语言模型和基于图形的注释中评估了嵌入.
主要成果:
- 生物知情的伪序列显著超过随机基线.
- 30-35个残留物的伪序列产生了最佳的预测性能.
- 结构和进化多样性伪序列的表现类似,表明重叠的残留物的重要性.
- 在ESM-2和注释嵌入中,比随机的伪序列有所改善,但比精选的伪序列没有改善.
结论:
- 策划的伪序列目前是新抗原预测模型中最有效的MHC表示.
- 虽然替代编码显示出希望,但它们尚未取代残留级序列信息的预测能力.
相关概念视频
Signal Sequences and Sorting Receptors
14.3K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
14.3K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
Maxam-Gilbert Sequencing
12.5K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
12.5K
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Sanger Sequencing
772.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.7K


