Related Experiment Video
Updated: Feb 8, 2026

T and B Cell Receptor Immune Repertoire Analysis using Next-generation Sequencing
Published on: January 12, 2021
Predicting the spectrum of TCR repertoire sharing with a data-driven model of recombination
Yuval Elhanati1, Zachary Sethna1, Curtis G Callan1
1Joseph Henry Laboratories, Princeton University, Princeton, NJ, USA.
Despite the extreme diversity of T-cell repertoires, many identical T-cell receptor (TCR) sequences are found in a large number of individual mice and humans. These widely shared sequences, often referred to as "public," have been suggested to be over-represented due to their potential immune functionality or their ease of generation by V(D)J recombination. Here, we show that even for large cohorts, the observed degree of sharing of TCR sequences between individuals is well predicted by a model accounting for the known quantitative statistical biases in the generation process, together with a simple model of thymic selection. Whether a sequence is shared by many individuals is predicted to depend on the number of queried individuals and the sampling depth, as well as on the sequence itself, in agreement with the data. We introduce the degree of publicness conditional on the queried cohort size and the size of the sampled repertoires. Based on these observations, we propose a public/private sequence classifier, "PUBLIC" (Public Universal Binary Likelihood Inference Classifier), based on the generation probability, which performs very well even for small cohort sizes.
Despite the extreme diversity of T-cell repertoires, many identical T-cell receptor (TCR) sequences are found in a large number of individual mice and humans. These widely shared sequences, often referred to as "public," have been suggested to be over-represented due to their potential immune functionality or their ease of generation by V(D)J recombination. Here, we show that even for large cohorts, the observed degree of sharing of TCR sequences between individuals is well predicted by a model accounting for the known quantitative statistical biases in the generation process, together with a simple model of thymic selection. Whether a sequence is shared by many individuals is predicted to depend on the number of queried individuals and the sampling depth, as well as on the sequence itself, in agreement with the data. We introduce the degree of publicness conditional on the queried cohort size and the size of the sampled repertoires. Based on these observations, we propose a public/private sequence classifier, "PUBLIC" (Public Universal Binary Likelihood Inference Classifier), based on the generation probability, which performs very well even for small cohort sizes.
Related Concept Videos
The Electromagnetic Spectrum
The Electromagnetic Spectrum
IR Spectrum
Transmittance is defined as the ratio of the radiant power passing through a sample to that from the radiation's source. Multiplying the transmittance by 100 gives the percent transmittance (%T), which varies between 100% (no absorption) and 0%...
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
Recombinant DNA
Predicting Molecular Geometry

