Assessing data size requirements for training generalizable sequence-based TCR specificity models via pan-allelic

Antoine Delaunay1, Miles McGibbon2, Bachir Djermani3

  • 1InstaDeep Ltd, 5 Merchant Sq, London, UK. a.delaunay@instadeep.com.

Scientific Reports
|November 27, 2025
PubMed

Insights

Developing personalized T cell receptor (TCR) therapies requires predicting TCR-neoepitope binding. Current machine learning models need millions of data points to generalize, far exceeding available data for effective oncology treatments.

Area of Science:

  • Immunology
  • Computational Biology
  • Oncology

Background:

  • Personalized T cell receptor (TCR) therapies for cancer rely on identifying TCRs that bind patient-specific neoepitopes.
  • The vast diversity of TCR and neoepitope repertoires presents a significant challenge for developing generalizable predictive models.

Purpose of the Study:

  • To estimate the required training data size for a machine learning model predicting TCR-pMHC specificity.
  • To assess the generalizability of existing models to unseen TCRs and epitopes.

Main Methods:

  • Validated existing models' limited generalizability to single-residue epitope dissimilarities.
  • Imputed the point-mutation ligandome across 34 human MHC alleles, representing it as a graph.
  • Determined the dominating set of the graph to estimate data requirements.

Main Results:

  • Published TCR specificity models fail to generalize beyond minimal epitope sequence variations.
  • An estimated 1 to 100 million unique epitopes are needed for a generalizable sequence-based TCR specificity prediction model.
  • This requirement is approximately 1000 times larger than current publicly available data.

Conclusions:

  • Current datasets are insufficient for training robust, generalizable TCR specificity prediction models for personalized cancer therapy.
  • Significant expansion of training data is necessary to advance the field of TCR-based oncology treatments.