Related Experiment Video
Updated: Aug 25, 2026

Use of Single Chain MHC Technology to Investigate Co-agonism in Human CD8+ T Cell Activation
Published on: February 28, 2019
Mitigating Goodhart's law in epitope-conditioned TCR generation using plug-and-play reward designs
Pengfei Zhang1,2, Xiaoyi He1,2, Fredo Guan1,2
1School of Computing Augmented Intelligence, Arizona State University, Tempe, AZ 85281, United States.
Motivation:
Epitope-conditioned T cell receptor (TCR) generation extends protein language modeling to the design of therapeutically relevant receptors. Reinforcement learning (RL) post-training with surrogate binding predictors can improve generation controllability, but it is vulnerable to Goodhart's Law: optimizing an imperfect surrogate reward can lead to reward inflation, distributional drift, and biologically implausible or nonspecific sequences. We investigate whether reward hacking can be mitigated through improved reward formulations without modifying the generator architecture or training pipeline and without requiring additional training data.
Results:
We introduce a plug-and-play reward-design framework for RL-based TCR generation that combines heuristic biological priors, model ensembling, and binding-specificity objectives based on max-margin and contrastive formulations. These components suppress degenerate sequences, reduce model-specific biases, and discourage cross-epitope binding. During RL fine-tuning, the proposed rewards stabilize optimization, limit surrogate-reward inflation, preserve canonical CDR3β sequence patterns and repertoire diversity, and maintain closer alignment with experimentally validated TCR-binder distributions. In evaluations on unseen epitopes, the specificity-aware reward formulations provide the strongest overall performance, improving diversity and ground-truth distributional alignment while retaining biological authenticity and predicted target binding. These findings demonstrate that Goodhart-resistant reward design improves the reliability, controllability, and generalization of epitope-conditioned TCR generation.
Availability And Implementation:
Code and models are available in a public repository (https://github.com/Lee-CBG/TCRRobustRewardDesign).
Related Concept Videos
T Cell Activation and Clonal Selection
Naive T cells that have not yet encountered an antigen express two primary CD...
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...

