通过大型语言模型对细胞类型和基因组注释进行基准测试,使用AnnDictionary.
George Crowley1, , Stephen R Quake2,3,4
1Department of Bioengineering, Stanford University, Stanford, California, USA.
Nature communications
|October 29, 2025
概括
一个开源软件包AnnDictionary允许使用大型语言模型 (LLM) 对anddata进行并行分析. 它对细胞类型和基因组注释的LLM进行了基准测试,发现主要细胞类型的准确性超过80%.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 人工智能的人工智能
背景情况:
- 单细胞RNA测序 (scRNA-seq) 可以产生大量的数据和数据集.
- 数据的自动化分析,特别是细胞类型的注释,至关重要.
- 大型语言模型 (LLM) 显示了生物数据分析的潜力.
研究的目的:
- 介绍AnnDictionary,这是一个开源软件包,用于使用LLMs进行并行和数据分析.
- 基准主要的LLM用于新的细胞类型注释准确性.
- 评估功能基因组注释中的LLM表现.
主要方法:
- 开发了AnnDictionary,集成LangChain和AnnData,支持多个LLM提供商.
- 实现了多线程优化,以高效地分析大数据.
- 进行了基准研究,将LLM注释与手册注释和基因组数据库进行比较.
主要成果:
- 细胞类型注释中的LLM性能因模型大小而异,主要细胞类型的高度一致 (>80-90%).
- 跨LLM协议也与模型大小相关.
- 克劳德 3.5 索内特在功能性基因组注释方面取得了高准确性 (>80%).
结论:
- AnnDictionary使用LLMs促进了对anddata的高效并行分析.
- 在scRNA-seq数据中,LLM显示了准确的细胞类型和功能注释的巨大潜力.
- 持续的基准测试和排名表维护将跟踪LLM在这个领域的进步.
相关概念视频
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Cell Specific Gene Expression
5.4K
5.4K
Cell Specific Gene Expression
16.2K
Multicellular organisms contain a variety of structurally and functionally distinct cell types, but the DNA in all the cells originated from the same parent cells. The differences in the cells can be attributed to the differential gene expression. Liver cells, whose functions include detoxification of blood, production of bile to metabolize fats, and synthesis of proteins essential for metabolism, must express a specific set of genes to perform their functions. Gene expression also varies with...
16.2K
Cell Lines
10.0K
A cell line is a population of cells grown in vitro that can be subcultured over several generations. Normal cells cease to divide after a certain number of cell divisions, a process known as replicative senescence. This number, called the Hayflick limit, was conceptualized by Leonard Hayflick in 1961 when he observed that fetal cells grown in culture could only divide 40-60 times. This limit is due to the shortening of the telomeres during each round of cell division, preventing cell division...
10.0K
Genetic Lingo
113.7K
Overview
113.7K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K

