Related Experiment Video
Updated: Jul 4, 2025

09:32
Immunopeptidomics: Isolation of Mouse and Human MHC Class I- and II-Associated Peptides for Mass Spectrometry Analysis
Published on: October 15, 2021
12.1K
Evaluating NetMHCpan performance on non-European HLA alleles not present in training data
Thomas Karl Atkins1, Arnav Solanki1, George Vasmatzis2
1Department of Electrical and Computer Engineering, University of Minnesota, Minneapolis, MN, United States.
Frontiers in Immunology
|January 31, 2024
Summary
Investigating machine learning models for healthcare applications is crucial for equity. Despite training data bias against certain ethnicities, NetMHCpan-4.1 and NetMHCIIpan-4.0 showed no significant drop in prediction accuracy for underrepresented groups.
Area of Science:
- Immunoinformatics
- Computational Biology
- Machine Learning in Healthcare
Background:
- Bias in neural network training datasets can reduce prediction accuracy for underrepresented groups, impacting healthcare applications.
- Machine learning models like NetMHCpan-4.1 and NetMHCIIpan-4.0 predict antigen binding to MHC molecules, critical for adaptive immune responses.
- Previous studies indicated potential bias in these models towards hydrophobic peptides.
Purpose of the Study:
- To investigate the composition of training datasets for NetMHCpan-4.1 and NetMHCIIpan-4.0, focusing on ethnic representation.
- To evaluate the prediction accuracy of these models for alleles not included in their training datasets.
- To understand the relationship between training data disparities, HLA sequence diversity, and prediction performance.
Main Methods:
- Analysis of training dataset composition for ethnic bias (European Caucasian, Asian, Pacific Islander).
- Testing NetMHCpan-4.1 and NetMHCIIpan-4.0 performance on alleles absent from training data.
- Mapping HLA sequence space to assess training dataset diversity.
- Linking impactful residues in NetMHCpan predictions to structural features for specific alleles.
Main Results:
- The training datasets for NetMHCpan-4.1 and NetMHCIIpan-4.0 were found to be heavily biased against Asian and Pacific Islander individuals, favoring European Caucasians.
- Unexpectedly, no meaningful difference in prediction quality was observed for alleles not present in the training data, despite the identified ethnic disparities.
- Analysis revealed the sequence diversity of the training dataset and linked key prediction residues to structural features.
Conclusions:
- Ethnic bias in training data for MHC-binding prediction models does not necessarily translate to decreased prediction quality for underrepresented alleles.
- Further investigation into HLA sequence space and structural features is necessary to fully understand model behavior and potential biases.
- Ensuring equitable performance of machine learning models in healthcare requires careful consideration of training data composition and its impact on diverse populations.

