Related Experiment Video
Updated: Jul 5, 2026

Using SCOPE to Identify Potential Regulatory Motifs in Coregulated Genes
Published on: May 31, 2011
Causal intervention validation of gene regulatory signals in scGPT
1Department of Computer Science, University of Tuebingen, Sand 14, 72076 Tuebingen, Germany.
Objective:
Single-cell foundation models such as scGPT have been promoted as representations of gene regulation, but their advertised regulatory signal has been judged largely from attention weights, which are correlational. We ask a deliberately bounded question: whether direct interventions on scGPT gene tokens-causal with respect to the model's own computation-recover model-internal transcription-factor (TF)-target dependencies that align with curated references, whether that signal is robust, and whether it transfers to real perturbation responses. We separate these into two explicit validation axes and report both, including where the model-internal signal does not correspond to biological causality.
Methods:
On Tabula Sapiens kidney, lung, and immune subsets, plus an external Krasnow lung atlas and three CRISPR perturbation datasets (Adamson, Dixit, Shifrut), we ablate or swap TF token values inside scGPT and quantify changes in target-token readouts. Axis 1 (reference alignment): intervention scores are evaluated against curated TRRUST and DoRothEA references and against attention and coexpression baselines, with new robustness studies over cell count (120-500), positive/negative pair count, four negative-sampling designs, and four readout strategies (mean, L2, cosine, max). We trace component-level circuits via activation patching, and benchmark against twelve GRN inference methods including a neural-network baseline on a matched substrate, additionally giving the classical methods larger cell budgets to test the equal-information question. Axis 2 (perturbation transfer): scores are evaluated against CRISPR perturbation-derived edges under balanced labels and AUROC.
Results:
On Axis 1, lung shows reproducible enrichment that is stable across cell count and survives all four negative-sampling designs (permutation p improving from 0.07 to 0.03 as cells grow from 120 to 500); kidney enrichment is significant at small samples but does not survive scaling (permutation p rising to 0.16 at 500 cells); immune is at baseline. Richer readouts (L2, cosine) recover signal the scalar mean compresses away, particularly in kidney. On the matched benchmark scGPT leads classical and neural baselines at the shared 120-cell budget, but the lead narrows as the classical methods are given more cells, confirming this is a signal-efficiency result, not general superiority. On Axis 2, balanced evaluation of all three CRISPR datasets, including the TF-screen Dixit data, yields AUROC ≈ 0.50: the model-internal signal does not predict real perturbation responses.
Conclusion:
scGPT encodes tissue-conditional, intervention-sensitive regulatory structure that is aligned with literature-curated TF-target edges (robustly in lung) but is representational rather than biologically causal: it does not transfer to perturbation outcomes. The pipeline is a practical mechanistic-audit toolkit for biological foundation models, and the gap between reference alignment and perturbation transfer is a concrete cautionary result for using such models in regulatory inference.
Related Concept Videos
Global Regulatory Systems
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Cis-regulatory Sequences
Combinatorial Gene Control
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...

