Related Experiment Video
Updated: Aug 6, 2026

Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing
Published on: March 15, 2019
Machine Learning for TCR Repertoire Epitope Annotation and Pattern Discovery
Romi Vandoren1,2,3, Vincent Van Deuren1,2,3, Fabio Affaticati1,2,3
1Adrem Data Lab, Department of Computer Science, University of Antwerp, Antwerp, Belgium.
Abstract:
T cells are central to adaptive immunity, recognizing antigenic peptides, called epitopes, via the T cell receptor (TCR). The immense diversity and cross-reactivity of the TCR repertoire makes direct interpretation of antigen specificity from repertoire sequencing challenging. High-throughput sequencing enables large-scale profiling of TCRs but does not directly reveal their target epitopes, requiring computational approaches to bridge this gap. This review outlines two complementary strategies, bottom-up and top-down approaches, to annotate TCR specificity. Bottom-up methods predict TCR-epitope specificity from curated TCR-epitope databases, identifying recurring patterns through distance-based, feature-based, or deep learning models. While effective for well-characterized epitopes, they are limited by biased training data, absence of negative data, and weak generalization to unseen epitopes. Top-down approaches instead infer antigen-driven responses from repertoire-level signals such as sequence similarity, enrichment, and TCR convergence. These methods enable discovery of disease- or exposure-associated TCR signatures without prior epitope knowledge but are sensitive to technical noise and biological confounding. Both approaches are complementary as bottom-up provides mechanistic specificity, while top-down enables discovery in complex datasets. Their integration, alongside multimodal modeling and improved benchmarking, is key to advancing TCR-epitope annotation and understanding adaptive immune responses.

