Related Experiment Video
Updated: Jun 13, 2025

A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
A Gene Set Foundation Model Pre-Trained on a Massive Collection of Diverse Gene Sets
Daniel J B Clarke1, Giacomo B Marino1, Avi Ma'ayan1
1Department of Pharmacological Sciences, Mount Sinai Center for Bioinformatics, Icahn School of Medicine at Mount Sinai, One Gustave L. Levy Place, Box 1603, New York, NY 10029 USA.
Abstract:
Trained with large datasets, foundation models can capture complex patterns within these datasets to create embeddings that can be used for a variety of useful applications. Here we created a gene set foundation model that was trained on a massive collection of unlabeled gene sets from two databases: Rummagene and RummaGEO. Rummagene automatically extracts gene sets from supplemental tables of publications; and RummaGEO has gene sets automatically computed from comparing groups of samples from RNA-seq studies deposited into the gene expression omnibus. Several foundation model architectures and data sources for training were benchmarked in the task of predicting gene function. Such predictions were also compared to other state-of-the-art gene function prediction methods and models. One of the GSFM architectures achieves superior performance compared to all other methods and models. This model was used to systematically predict gene functions for all human gene. These predictions are served on gene pages that are accessible from https://gsfm.maayanlab.cloud.
More Related Videos
Related Concept Videos
Genomics
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
DNA Microarrays
Organization of Genes
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Genome Annotation and Assembly

