RummaGEO:从GEO中自动挖掘人类和老鼠基因组
Giacomo B Marino1, Daniel J B Clarke1, Eden Z Deng1
1Mount Sinai Center for Bioinformatics, Department of Pharmacological Sciences, Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York 10029, NY USA.
bioRxiv : the preprint server for biology
|April 22, 2024
概括
RummaGEO使基因表达总 (GEO) 存储库的全球数据级搜索成为可能. 这种新工具在众多人类和小鼠RNA-seq研究中促进基因表达签名搜索,帮助生物医学研究和假设生成.
科学领域:
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
- 计算生物学 计算生物学
背景情况:
- 基因表达总汇 (GEO) 是一个庞大的转录学数据库,但缺乏数据级搜索功能.
- 现有的GEO搜索方法集中在元数据上,限制了直接的数据探索和分析.
研究的目的:
- 开发RummaGEO,这是一个用于在GEO内全面搜索基因表达签名的Web服务器.
- 为了使人类和小鼠的RNA测序研究能够在数据层面上进行探索.
主要方法:
- 利用ARCHS4的统一对齐的GEO研究来离线识别样本条件.
- 计算差异表达特征,从已识别的研究中提取基因组.
- 将签名搜索,PubMed搜索和元数据搜索功能集成到一个Web服务器中.
主要成果:
- 鲁玛GEO包含来自23,395个GEO研究的135,264个人类和158,062个小鼠基因组.
- 该数据库支持全球分析和识别GEO数据中的统计模式.
- 网络服务器为假设生成提供了前所未有的资源.
结论:
- RummaGEO解决了在GEO存储库中进行数据级搜索的需求.
- 该工具通过有效探索基因表达数据来增强生物医学研究.
- 鲁玛GEO提供了一个有价值的资源,用于从现有的转录基因数据中生成新的假设.
相关概念视频
Mouse Models of Cancer Study
5.5K
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
5.5K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K


