Related Experiment Video
Updated: Aug 1, 2026

Measuring TCR-pMHC Binding In Situ using a FRET-based Microscopy Assay
Published on: October 30, 2015
Assessing data size requirements for training generalizable sequence-based TCR specificity models via pan-allelic
Antoine Delaunay1, Miles McGibbon2, Bachir Djermani3
1InstaDeep Ltd, 5 Merchant Sq, London, UK. a.delaunay@instadeep.com.
Abstract:
Rapid identification of T cell receptors (TCRs) that specifically bind patient-unique neoepitopes is a critical challenge for personalized TCR-based therapies in oncology. Due to enormous diversity of both TCR and neoepitope repertoires, a machine learning predictor of TCR-pMHC specificity for personalized therapy must generalize to TCRs and epitopes not seen in the training data. We estimate the necessary size of such training data. We first confirm that published models fail to generalize beyond a single-residue dissimilarity to the epitope training set distribution. We then impute the point-mutation ligandome across the 34 most prevalent human MHC alleles and represent it as a graph based on our established dissimilarity cutoff. By finding the dominating set of this graph, we estimate that between one and 100 million epitopes are required to train a generalizable sequence-based TCR specificity prediction model-1000 times the size of current public data.
Insights
Developing personalized T cell receptor (TCR) therapies requires predicting TCR-neoepitope binding. Current machine learning models need millions of data points to generalize, far exceeding available data for effective oncology treatments.
Area of Science:
- Immunology
- Computational Biology
- Oncology
Background:
- Personalized T cell receptor (TCR) therapies for cancer rely on identifying TCRs that bind patient-specific neoepitopes.
- The vast diversity of TCR and neoepitope repertoires presents a significant challenge for developing generalizable predictive models.
Purpose of the Study:
- To estimate the required training data size for a machine learning model predicting TCR-pMHC specificity.
- To assess the generalizability of existing models to unseen TCRs and epitopes.
Main Methods:
- Validated existing models' limited generalizability to single-residue epitope dissimilarities.
- Imputed the point-mutation ligandome across 34 human MHC alleles, representing it as a graph.
- Determined the dominating set of the graph to estimate data requirements.
Main Results:
- Published TCR specificity models fail to generalize beyond minimal epitope sequence variations.
- An estimated 1 to 100 million unique epitopes are needed for a generalizable sequence-based TCR specificity prediction model.
- This requirement is approximately 1000 times larger than current publicly available data.
Conclusions:
- Current datasets are insufficient for training robust, generalizable TCR specificity prediction models for personalized cancer therapy.
- Significant expansion of training data is necessary to advance the field of TCR-based oncology treatments.

