Related Experiment Video
Updated: May 28, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
mmContext: an open framework for multimodal contrastive learning of omics and text data
Jonatan Menger1,2, Sonia Maria Krissmer1,2, Clemens Kreutz1,2
1Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center, University of Freiburg, Freiburg im Breisgau, 79104, Germany.
We developed mmContext, a new framework for integrating omics and text data. This tool enables researchers to compare different text encoders for multimodal learning in biology.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning in Biology
Background:
- Multimodal approaches are crucial for integrating omics data with biological text.
- Existing frameworks lack standardization for comparing omics representations with various text encoders.
- A unified workflow is needed for systematic analysis of multimodal biological data.
Purpose of the Study:
- To introduce mmContext, a novel, accessible, and standardized framework for multimodal learning.
- To enable systematic comparison of omics data embeddings generated with different text encoders.
- To facilitate reproducible and accessible multimodal learning for omics-text integration.
Main Methods:
- Developed mmContext, a lightweight and extensible multimodal embedding framework using Sentence Transformers.
- Integrated omics data (numeric representation in AnnData.obsm) with text data using Hugging Face text encoders.
- Provided pipelines for training, evaluation, and data preparation for multimodal models.
Main Results:
- Trained and evaluated models for RNA-Seq and text integration.
- Demonstrated utility through zero-shot classification of cell types and diseases across four independent datasets.
- Showcased mmContext's capability in multimodal learning for biological data.
Conclusions:
- mmContext provides a standardized framework for multimodal omics-text integration.
- The framework enables reproducible research through open-source models, datasets, and tutorials.
- Facilitates advanced biological insights by combining diverse data types effectively.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Genomics
Tandem Mass Spectrometry
Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes
MicroRNAs