Related Experiment Video
Updated: Apr 25, 2026

Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
Discovery of candidate therapeutic targets with Geneformer
Yujie Zhang1,2,3, Madhavan S Venkatesh1,2,4, Christina V Theodoris5,6,7,8
1Gladstone Institute of Cardiovascular Disease, San Francisco, CA, USA.
Abstract:
Mapping the connections between genes enables the identification of networks disrupted in disease. The approach, however, requires large amounts of data, making the discovery of therapeutic targets difficult in settings with limited data. We recently developed a foundational artificial intelligence model, Geneformer, pretrained on a large-scale corpus of single-cell transcriptomes (initially ~30 million, now >100 million) to enable context-aware predictions in network biology with limited data. Here, we cover the methodology for using Geneformer through a combination of zero-shot inference, fine-tuning and in silico perturbation. The procedure includes the tokenization of raw gene expression counts into rank value encodings aligned with the model's pretrained vocabulary. Separability of relevant phenotypes in the pretrained embedding space is first assessed with zero-shot embeddings. Fine-tuning is then performed either with a single task, for example, disease prediction within a specific cell type, or with multiple tasks to jointly learn cross-informative features, such as cell types and disease states. Performance is evaluated with confusion matrices, macro F1 scores and embedding analysis. Subsequently, in silico perturbation simulates gene repression or activation and quantifies the shift in cell state embeddings, prioritizing candidate targets by statistical and biological metrics. The approach also supports perturbation using a quantized model to enhance computational efficiency. Outputs include predictive models fine-tuned for context-specific cell state representations and rank-ordered predictions of perturbations to induce each target state. The full pipeline typically completes in under 2 days on a standard GPU workstation and requires only moderate Python experience.
More Related Videos
09:33Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
09:37Defining Gene Functions in Tumorigenesis by Ex vivo Ablation of Floxed Alleles in Malignant Peripheral Nerve Sheath Tumor Cells
Published on: August 25, 2021
Related Concept Videos
Pharmacogenomics: Identification of New Drug Targets
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Gene Therapy
Gene Therapy
Drug Discovery: Overview
In-vitro Mutagenesis