Related Experiment Video
Updated: Oct 7, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
BioMamba: Domain-Adaptive Biomedical Language Models
Ling Yue1, Mingzhi Zhu1, Sixue Xing1
1Department of Computer Science, Rensselaer Polytechnic Institute, Troy, NY, USA.
Abstract:
Background: Biomedical language models should improve performance on biomedical text while retaining general-language-modeling fluency. For Mamba-based models, this trade-off has not been systematically studied across biomedical literature and clinical text. Methods: We developed BioMamba, a family of biomedical Mamba2 models at 5 scales (130M, 370M, 780M, 1.3B, and 2.7B) obtained by continued pretraining of released public Mamba2 checkpoints on a balanced 80%/10%/10% mixture of PubMed abstracts, the Colossal Clean Crawled Corpus (C4), and Wikipedia. The architecture is unchanged from Mamba2; the contribution is the adaptation recipe and the accompanying open-weight checkpoints. We evaluated internal language modeling on fixed held-out sets, out-of-domain multiple-choice benchmarks, and 3 downstream tasks across multiple model scales: clinical note completion and discharge summary generation on MIMIC-IV-Note, and biomedical yes/no question answering on Biomedical Semantic Indexing and Question Answering (BioASQ) and PubMed Question Answering (PubMedQA). Results: Across 5 scales, BioMamba consistently lowered PubMed perplexity, improved Wikipedia-style held-out perplexity by 1.46 to 4.72 PPL, and left C4 perplexity essentially unchanged ( ). On 6 out-of-domain multiple-choice benchmarks, BioMamba stayed within ±3 pp of Mamba2 with no systematic regression. After supervised fine-tuning, BioMamba+SFT matched or exceeded Mamba2+SFT on MIMIC-IV note completion and discharge summary generation at every evaluated scale (paired-bootstrap at 2.7B) and improved PubMedQA at every scale. The strongest model (BioMamba-2.7B) reached a PubMed perplexity of 5.28 and accuracies of 90.24% and 73.00% on BioASQ and PubMedQA, respectively. Conclusions: A balanced domain-adaptive continued pretraining recipe strengthens Mamba2 language models on biomedical literature and clinical text while preserving general-language-modeling fluency.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Genome Annotation and Assembly
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Bioremediation
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Applications Of NMR In Biology
The...