细胞2句子:教大语言模型生物学的语言
Daniel Levine1, Syed Asad Rizvi1, Sacha Lévy1
1Department of Computer Science, Yale University, New Haven, CT, USA.
bioRxiv : the preprint server for biology
|November 18, 2024
概括
Cell2Sentence (C2S) 将基因表达数据转化为"细胞句子",使大型语言模型能够理解单细胞生物学. 这种方法促进了细胞生成和准确的细胞类型注释,用于各种生物应用.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 单细胞转录组产生高维基基因表达数据.
- 集成先进的计算方法,如自然语言处理 (NLP) 可以解锁新的见解.
- 现有的NLP模型需要对生物数据进行调整.
研究的目的:
- 引入Cell2Sentence (C2S),一种用于将大型语言模型 (LLM) 适应单细胞转录学的新方法.
- 证明C2S在使LLM能够执行生物任务方面的实用性.
- 为了弥合NLP和单细胞生物学之间的差距.
主要方法:
- 将基因表达数据转化为"细胞句子".
- 使用单元格句子微调预先训练的LLM (例如,GPT-2).
- 在诸如细胞生成和细胞类型注释等任务上评估微调模型.
主要成果:
- 精心调整的GPT-2模型可以根据细胞类型输入生成生物有效的细胞.
- 这些模型准确地从细胞句子中预测细胞类型.
- 与C2S微调的LLM获得了对单细胞生物学的重要理解.
结论:
- Cell2Sentence (C2S) 提供了一个灵活的框架,可以将NLP与转录学集成在一起.
- C2S使LLM能够执行复杂的生物任务,包括细胞生成和注释.
- 这种方法利用现有的NLP模型和库用于广泛的生物应用.
更多相关视频
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
3.9K
10:41Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
9.2K
相关概念视频
Proteins: From Genes to Degradation
12.0K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
Transcription is the synthesis of RNA...
12.0K
Improving Translational Accuracy
9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Genome Size and the Evolution of New Genes
2.4K
2.4K
Calmodulin-dependent Signaling
5.1K
Calmodulin (CaM) is a calcium-binding protein in eukaryotes that controls various calcium-regulated cellular processes. It has four calcium-binding sites that bind calcium to form the calcium-calmodulin ( Ca2+-CaM) complex. GPCR stimulation increases the calcium levels in the cells that bind to CaM and induces a conformational change.
The Ca2+-CaM complex does not have enzymatic activity by itself. Instead, the complex binds downstream target proteins, including membrane proteins or enzymes,...
The Ca2+-CaM complex does not have enzymatic activity by itself. Instead, the complex binds downstream target proteins, including membrane proteins or enzymes,...
5.1K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
