Related Experiment Video
Updated: May 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Evaluating Expert Specialization in Mixture-of-Experts Antibody Language Models
Sarah M Burbach1, Simone Spandau1, Jonathan Hurtado1
1Department of Immunology and Microbiology, The Scripps Research Institute, USA.
Biorxiv : the Preprint Server for Biology
|May 4, 2026
Summary
Antibody language models (AbLMs) using a sparse Mixture-of-Experts (MoE) architecture improve learning of diverse antibody regions. This approach outperforms dense models, enhancing antibody sequence modeling capabilities.
Area of Science:
- Computational biology
- Bioinformatics
- Artificial intelligence in medicine
Background:
- Antibody language models (AbLMs) excel at learning antibody features but struggle with highly diverse, non-templated regions.
- Current AbLMs utilize dense architectures where all parameters attend to every amino acid token.
- Mixture-of-Experts (MoE) architectures, common in natural language processing, offer potential for specialized parameter use but are less explored in biological modeling.
Purpose of the Study:
- To investigate the efficacy of a sparse Mixture-of-Experts (MoE) architecture for Antibody Language Models (AbLMs).
- To adapt and optimize MoE routing strategies for antibody sequence data, particularly focusing on CDRH3 regions.
- To evaluate if an MoE-based AbLM can outperform dense models in learning antibody features.
Main Methods:
- Assessed existing MoE routing strategies, comparing token-choice and expert-choice routing for AbLMs.
- Optimized the token-choice router to minimize padding token routing, enabling pre-training with variable sequence lengths.
- Developed and trained a large-scale baseline antibody language model with a Top-2 MoE architecture (BALM-MoE) on diverse antibody sequences.
Main Results:
- Token-choice routing strategies demonstrated superior performance over expert-choice routing in AbLMs, likely due to specialization in CDRH3 residues.
- The optimized token-choice router effectively handled variable sequence lengths by minimizing padding token engagement.
- The BALM-MoE model, with a Top-2 MoE architecture, outperformed its dense counterpart with an equivalent number of active parameters.
Conclusions:
- Sparse MoE architectures are beneficial for AbLMs, enabling better learning of antibody sequence diversity compared to dense models.
- Optimized MoE routing strategies enhance the applicability of AbLMs for biological sequence modeling.
- MoE-based AbLMs represent a promising advancement for antibody design and analysis.
