Systematically benchmarking peptide-MHC binding predictors: From synthetic to naturally processed epitopes

Weilong Zhao1, Xinwei Sher1

  • 1Global Research IT, Merck & Co., Inc., Boston, MA, United States of America.

A number of machine learning-based predictors have been developed for identifying immunogenic T-cell epitopes based on major histocompatibility complex (MHC) class I and II binding affinities. Rationally selecting the most appropriate tool has been complicated by the evolving training data and machine learning methods. Despite the recent advances made in generating high-quality MHC-eluted, naturally processed ligandome, the reliability of new predictors on these epitopes has yet to be evaluated. This study reports the latest benchmarking on an extensive set of MHC-binding predictors by using newly available, untested data of both synthetic and naturally processed epitopes. 32 human leukocyte antigen (HLA) class I and 24 HLA class II alleles are included in the blind test set. Artificial neural network (ANN)-based approaches demonstrated better performance than regression-based machine learning and structural modeling. Among the 18 predictors benchmarked, ANN-based mhcflurry and nn_align perform the best for MHC class I 9-mer and class II 15-mer predictions, respectively, on binding/non-binding classification (Area Under Curves = 0.911). NetMHCpan4 also demonstrated comparable predictive power. Our customization of mhcflurry to a pan-HLA predictor has achieved similar accuracy to NetMHCpan. The overall accuracy of these methods are comparable between 9-mer and 10-mer testing data. However, the top methods deliver low correlations between the predicted versus the experimental affinities for strong MHC binders. When used on naturally processed MHC-ligands, tools that have been trained on elution data (NetMHCpan4 and MixMHCpred) shows better accuracy than pure binding affinity predictor. The variability of false prediction rate is considerable among HLA types and datasets. Finally, structure-based predictor of Rosetta FlexPepDock is less optimal compared to the machine learning approaches. With our benchmarking of MHC-binding and MHC-elution predictors using a comprehensive metrics, a unbiased view for establishing best practice of T-cell epitope predictions is presented, facilitating future development of methods in immunogenomics.

Related Concept Videos

Peptide Bonds02:43

Peptide Bonds

A peptide bond covalently attaches amino acids through a dehydration reaction. One amino acid's carboxyl group and another amino acid's amino group combine, releasing a water molecule. The resulting bond is the peptide bond. The products that such linkages form are peptides. As more amino acids join this growing chain, the resulting chain is a polypeptide. Each polypeptide has a free amino group at one end. This end has the N-terminal, or the amino-terminal, and the other end has a free...
83.0K
What is Natural Selection?01:32

What is Natural Selection?

Natural selection is an evolutionary process in which individuals with survival-promoting traits reproduce at higher rates. These favorable traits become more common within a population or species. Naturally selected traits initially arise via random genetic mutations. In order for selection to occur, there must be variation within a population, the trait controlling the variation must be heritable, and there must be an evolutionary advantage for variation in the trait.
129.2K
The Equilibrium Binding Constant and Binding Strength02:18

The Equilibrium Binding Constant and Binding Strength

The equilibrium binding constant (Kb) quantifies the strength of a protein-ligand interaction. Kb can be calculated as follows when the reaction is at equilibrium:
15.1K
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.2K
Nature and Nurture01:10

Nature and Nurture

Many human characteristics, like height, are shaped by both nature—in other words, by our genes—and by nurture, or our environment. For example, chronic stress during childhood inhibits the production of growth hormones and consequently reduces bone growth and height. Scientists estimate that 70-90% of variation in height is due to genetic differences among individuals, and 10-30% of variation in height is due to differences in the environments that individuals experience,...
22.4K
Random and Systematic Errors01:20

Random and Systematic Errors

Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.9K