Nucleotide context models outperform protein language models for predicting antibody affinity maturation

Mackenzie M Johnson1, Kevin Sung1, Hugh K Haddox1

  • 1Computational Biology Program, Fred Hutchinson Cancer Center, Seattle, Washington, United States of America.

PubMed

Antibodies play a crucial role in adaptive immunity. They develop as B cell receptors (BCRs): membrane-bound forms of antibodies that are expressed on the surfaces of B cells. BCRs are refined through affinity maturation, a process of somatic hypermutation (SHM) and natural selection, to improve binding to an antigen. Computational models of affinity maturation have developed from two main perspectives: molecular evolution and language modeling. The molecular evolution perspective focuses on nucleotide sequence context to describe mutation and selection; the language modeling perspective involves learning patterns from large data sets of protein sequences. In this paper, we compared models from both perspectives on their ability to predict the course of antibody affinity maturation along phylogenetic trees of BCR sequences. This included models of SHM, models of SHM combined with an estimate of selection, and protein language models. We evaluated these models for large human BCR repertoire data sets, as well as an antigen-specific mouse experiment with a pre-rearranged cognate naive antibody. We demonstrated that precise modeling of SHM, which requires the nucleotide context, provides a substantial amount of predictive power for predicting the course of affinity maturation. Notably, a simple nucleotide-based convolutional neural network modeling SHM outperformed state-of-the-art protein language models, including one trained exclusively on antibody sequences. Furthermore, incorporating estimates of selection based on a custom deep mutational scanning experiment brought only modest improvement in predictive power. To support further research, we introduce EPAM (Evaluating Predictions of Affinity Maturation), a benchmarking framework to integrate evolutionary principles with advances in language modeling, offering a road map for understanding antibody evolution and improving predictive models.

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Antibody Structure01:10

Antibody Structure

Overview
Antibodies, also known as immunoglobulins (Ig), are essential players of the adaptive immune system. These antigen-binding proteins are produced by B cells and make up 20 percent of the total blood plasma by weight. In mammals, antibodies fall into five different classes, which each elicits a different biological response upon antigen binding.
The Y-Shaped Structure of Antibodies Consists of Four Polypeptide Chains
Antibodies consist of four polypeptide chains: two identical heavy...
65.2K
Antibody Structure and Classes01:25

Antibody Structure and Classes

Antibodies, also known as immunoglobulins, are produced by B cells in response to foreign substances, such as bacteria and viruses. These proteins are critical for recognizing and neutralizing these substances, protecting the body from potential harm.
The basic structure of an antibody consists of four protein chains: two identical heavy chains and two identical light chains. These chains are held together by disulfide bonds and other non-covalent interactions, forming a Y-shaped structure.
8.1K
Physiological Pharmacokinetic Models: Assumption with Protein Binding01:13

Physiological Pharmacokinetic Models: Assumption with Protein Binding

Physiological models with protein binding in pharmacokinetics offer a sophisticated approach to understanding drug disposition. These models consider drug-protein interactions, enabling them to effectively predict drug concentrations in different organs and tissues. This precision aids in accurate drug dosing, providing a significant advantage over conventional models. A key process within these models is equilibration, which ensures that drug concentrations achieve a steady state within the...
201