Related Experiment Video
Updated: May 17, 2025

09:35
A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
17.7K
Using semantic search to find publicly available gene-expression datasets
Grace S Brown1, James Wengler1,2, Aaron Joyce S Fabelico1
1Department of Biology, Brigham Young University, Provo, Utah, USA.
Biorxiv : the Preprint Server for Biology
|March 31, 2025
Summary
Language models can enhance the discovery of relevant scientific datasets by summarizing descriptions into embeddings. This approach aids researchers in finding similar data for reuse and validation, improving upon existing search methods.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Science
Background:
- Vast numbers of high-throughput molecular datasets are publicly available in repositories like Gene Expression Omnibus (GEO).
- Reusing these datasets is crucial for validating findings and exploring new research questions.
- Discovering relevant datasets is challenging due to sheer volume, inconsistent descriptions, and lack of semantic annotations, hindering FAIR data principles.
Purpose of the Study:
- To evaluate the effectiveness of language models in improving dataset discovery within the Gene Expression Omnibus (GEO).
- To assess if language model-generated embeddings can identify relevant datasets more efficiently than traditional search methods.
Main Methods:
- Utilized 30 language models to generate numerical representations (embeddings) of dataset descriptions from GEO.
- Focused on six human medical conditions, using datasets previously curated by humans.
- Compared the performance of language model-based similarity searches against GEO's built-in search engine.
Main Results:
- Language models, particularly those trained on general corpora using contrastive learning with large embeddings, often outperformed GEO's search engine in identifying relevant datasets.
- The effectiveness varied, indicating that this approach is promising but not universally superior.
- Identified specific model characteristics that correlate with better performance in dataset discovery.
Conclusions:
- Language models show significant potential to improve the discovery of scientific datasets, complementing existing search tools.
- This approach can aid researchers in efficiently finding and reusing valuable molecular data.
- Further development and integration of language models could streamline data discovery and enhance scientific reproducibility.
Related Concept Videos
DNA Microarrays
17.1K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.1K
Reporter Genes
11.1K
Reporter genes are a type of protein-coding gene that are often tagged to a gene of interest. Once inside a target cell, reporter genes usually produce visually identifiable characteristics like fluorescence and luminescence when expressed along with the gene of interest. Thus, reporter genes “report” the presence or absence of genes of interest in an organism, determine the gene expression pattern, or track the physical location of a DNA segment or protein in the cell.
11.1K
Cell Specific Gene Expression
4.5K
4.5K

