ネオアンチゲン予測において擬似配列は重要か?
bioRxiv : the preprint server for biology
|December 22, 2025
まとめ
生物学的に情報に基づいた擬似配列は、個別化がんワクチンにおける正確なネオアンチゲン予測に不可欠である。特に30~35残基のキュレーションされた擬似配列は、T細胞応答を予測するための代替エンコーディングよりも依然として優れている。
科学分野:
- 免疫情報学
- 計算生物学
- がん免疫学
背景:
- 個別化がんワクチンは、T細胞応答を誘発するネオアンチゲンを予測することに依存しています。
- 現在の方法ではMHCクラスI擬似配列がしばしば使用されますが、その最適な定義とエンコーディングは不明確です。
研究 の 目的:
- ネオアンチゲン予測のためのさまざまな主要組織適合遺伝子複合体(MHC)クラスI表現を体系的に評価すること。
- 構造ベース、進化的、ランダム、およびさまざまな長さを含む擬似配列戦略と、タンパク質言語モデルおよびグラフアノテーションからの埋め込みを比較すること。
主な方法:
- BigMHC ELフレームワークを使用して、さまざまなMHCアレル表現を評価しました。
- 構造、進化的多様性などの生物学的に情報に基づいた擬似配列のパフォーマンスを、ランダムベースラインおよびさまざまな長さと比較しました。
- ESM-2タンパク質言語モデルおよびグラフベースアノテーションからの埋め込みのパフォーマンスを評価しました。
主要な成果:
- 生物学的に情報に基づいた擬似配列は、ランダムベースラインを大幅に上回りました。
- 30~35残基の擬似配列が最適な予測パフォーマンスをもたらしました。
- 構造と進化的多様性の擬似配列は同様のパフォーマンスを示し、残基の重要性が重複していることを示唆しています。
- ESM-2およびアノテーション埋め込みは、ランダムよりも改善を示しましたが、キュレーションされた擬似配列よりも改善しませんでした。
結論:
- キュレーションされた擬似配列は、現在、ネオアンチゲン予測モデルにとって最も効果的なMHC表現です。
- 代替エンコーディングは有望ですが、残基レベルの配列情報の予測力をまだ置き換えるものではありません。
関連する概念動画
Signal Sequences and Sorting Receptors
14.3K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
14.3K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
Maxam-Gilbert Sequencing
12.5K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
12.5K
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Sanger Sequencing
772.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.7K


