Benchmarking foundation cell models for post-perturbation RNA-seq prediction.
Gerold Csendes1, Gema Sanz1, Kristóf Z Szalay1
1Turbine Ltd., Budapest, Hungary.
BMC Genomics
|April 24, 2025
Summary
Predicting cellular responses to perturbations is crucial. Current foundation cell models like scGPT and scFoundation underperform simple baselines, indicating issues with benchmarking and datasets for gene expression prediction.
Area of Science:
- Computational Biology
- Genomics
- Systems Biology
Background:
- Accurate prediction of cellular responses to perturbations is vital for understanding cell behavior in health and disease.
- Foundation cell models pre-trained on large-scale single-cell gene expression data are state-of-the-art for predicting post-perturbation profiles.
- However, robust benchmarking of these models remains a significant challenge due to data limitations.
Purpose of the Study:
- To benchmark the performance of recently developed foundation cell models (scGPT, scFoundation) against baseline models for predicting gene expression after cellular perturbations.
- To identify limitations in current benchmarking methodologies and benchmark datasets for evaluating post-perturbation gene expression prediction models.
Main Methods:
- Comparative benchmarking of scGPT and scFoundation against various baseline models, including simple statistical methods and machine learning models incorporating biological features.
- Evaluation of model performance on existing Perturb-Seq benchmark datasets.
Main Results:
- Surprisingly, the simplest baseline model (mean of training examples) outperformed both scGPT and scFoundation.
- Machine learning models incorporating biologically meaningful features significantly outperformed scGPT.
- Perturb-Seq benchmark datasets were found to have low perturbation-specific variance, rendering them suboptimal for model evaluation.
Conclusions:
- Current foundation cell models may not offer significant advantages over simpler methods for post-perturbation gene expression prediction.
- Existing benchmark datasets and evaluation strategies are insufficient for accurately assessing the performance of these advanced models.
- Further research is needed to develop more effective benchmarking approaches and datasets for robust evaluation of cellular response prediction models.
Related Concept Videos
RNA-seq
9.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.7K
Ribosome Profiling
3.4K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.4K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Nonsense-mediated mRNA Decay
10.3K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.3K


