Continued domain-specific pre-training of protein language models for pMHC-I binding prediction
Sergio E Mares1, Ariel Espinoza Weinberger2, Nilah M Ioannidis3
1Center for Computational Biology, University of California, Berkeley, CA, USA.
Abstract:
Predicting peptide-major histocompatibility complex I (pMHC-I) binding affinity remains challenging due to extreme allelic diversity (~30,000 HLA alleles), severe data scarcity for most alleles, and noisy experimental measurements. Current methods particularly struggle with underrepresented alleles and quantitative binding prediction. We test whether domain-specific continued pre-training of protein language models is beneficial for their application to pMHC-I binding affinity prediction. Starting from ESM Cambrian (300M parameters), we perform masked-language modeling (MLM)-based continued pre-training on HLA-associated peptides (epitopes), testing two input formats: epitope sequences alone versus epitopes concatenated with HLA heavy chain sequences. We then fine-tune for functional IC50 binding affinity prediction using high-quality quantitative data.
Key Results:
Continued pre-training on epitope sequences improves model performance in data-scarce settings, with pretrained models achieving higher Spearman correlation than non-pretrained baselines at 250 training peptides before converging at larger training sizes. After continued pre-training and fine-tuning, our resulting model (ESMCBA) achieves a Spearman correlation of 0.61 on a held-out test set for predicting binding affinity across 24 common HLA alleles, competing with NetMHCpan (0.56), MHCflurry (0.49), and other state-of-the-art predictors.
Limitations:
The benefits of continued pre-training are most pronounced at moderate data availability (250-1500 peptides), with diminishing returns as training data increases beyond 3000 peptides, where pretrained and non-pretrained models converge to similar performance. Additionally, the method requires substantial computational resources and performance remains fundamentally limited by the inherent noise and experimental heterogeneity in binding affinity measurements from diverse assay protocols.
Impact:
This work demonstrates that domain-specific continued pre-training improves data efficiency for protein language models on specialized prediction tasks, particularly in low-resource settings. The finding that epitope-specific pretraining provides the largest performance gains at 250-1500 training examples has important implications for neoantigen vaccine prioritization, where many clinically relevant HLA alleles lack extensive binding data. More broadly, this study establishes a methodological framework for applying continued pre-training to other specialized biological prediction tasks where task-specific data is scarce but related unlabeled sequences are abundant.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Organization
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...


