Related Experiment Video
Updated: Jun 3, 2025

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
634
SampleExplorer: using language models to discover relevant transcriptome data
Wee Loong Chin1,2,3, Timo Lassmann3
1National Centre for Asbestos Related Diseases, QEII Medical Centre, Nedlands, WA 6009, Australia.
Bioinformatics (Oxford, England)
|January 9, 2025
Summary
SampleExplorer enhances biomedical research by enabling efficient discovery of relevant transcriptomic datasets. This tool leverages RNA-sequencing metadata and a language model to improve data retrieval for replication and verification studies.
Area of Science:
- Biomedical Research
- Bioinformatics
- Computational Biology
Background:
- Transcriptomics, particularly RNA-sequencing (RNA-seq), is a cornerstone of modern biomedical research.
- Large public repositories contain vast amounts of RNA-seq data with associated metadata, crucial for study understanding and replication.
- Existing metadata is underutilized for discovering relevant datasets, hindering efficient data reuse.
Purpose of the Study:
- To introduce SampleExplorer, a novel tool designed to facilitate the discovery of relevant transcriptomic datasets.
- To enable researchers to search for data using both text-based queries and gene set information.
- To improve the identification and accessibility of gene expression datasets within large public repositories.
Main Methods:
- SampleExplorer embeds sample metadata and utilizes a transformer-based language model for data retrieval.
- The tool was benchmarked using the ARCHS4 database to assess its effectiveness.
- Implementation details and algorithmic descriptions are available in supplementary materials.
Main Results:
- SampleExplorer effectively retrieves biologically relevant samples from large-scale transcriptomic data.
- The tool provides an efficient method for discovering relevant gene expression datasets.
- It enhances sample and dataset identification across diverse experimental contexts.
Conclusions:
- SampleExplorer offers a powerful solution for leveraging existing transcriptomic data.
- The tool supports replication and verification studies by improving data discoverability.
- It represents a significant advancement in utilizing RNA-seq metadata for research.
Related Concept Videos
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K
Improving Translational Accuracy
2.5K
2.5K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
lncRNA - Long Non-coding RNAs
2.8K
2.8K

