SummArIzeR:大規模言語モデルによるクロスデータベースエンリッチメント結果のクラスタリングと注釈の簡略化
Marie Brinkmann1, Michael Bonelli1, Anela Tosevska1
1Division of Rheumatology, Department of Internal Medicine III, Medical University of Vienna, Austria.
Bioinformatics (Oxford, England)
|March 2, 2026
まとめ
SummArIzeRは、クラスタリングおよび注釈付けエンリッチメント結果によって生物学的データ解釈を簡略化する新しいRパッケージです。高速で偏りのない注釈付けに大規模言語モデルを使用し、条件間での比較を改善します。
科学分野:
- バイオインフォマティクス
- 計算生物学
- ゲノミクス
背景:
- 複数のデータベースにわたるエンリッチメント分析は、冗長な用語につながり、生物学的データ解釈を複雑にします。
- エンリッチメント分析における重複する用語は、条件間での高速で直感的な解釈と比較を妨げます。
- 既存のツールには、複雑なエンリッチメント結果のクラスタリングと注釈付けのための効率的な方法が欠けています。
研究 の 目的:
- 複数のデータベースにわたるエンリッチメント結果のクラスタリングと注釈付けのためのRパッケージ、SummArIzeRを開発すること。
- 複数の条件にわたる生物学的データの高速で直感的な解釈と比較を可能にすること。
- 大規模言語モデルを使用したエンリッチメントクラスターの注釈付けを容易にすること。
主な方法:
- SummArIzeRは、共有遺伝子に基づいてエンリッチメント結果をクラスタリングします。
- クラスターごとにプールされたp値が計算されます。
- クラスター注釈には大規模言語モデルが利用されます。
- このパッケージは、結果の解釈しやすい視覚化を提供します。
主要な成果:
- SummArIzeRは、大規模言語モデルによって強化された、偏りのない高速なクラスター注釈を提供します。
- このパッケージは、手動キュレーションに匹敵するクラスタリングを実現します。
- SummArIzeRは、共有遺伝子に基づいてエンリッチメント結果の優れたグルーピングを提供します。
結論:
- SummArIzeRは、生物学的エンリッチメント分析の解釈を強化します。
- このRパッケージは、複雑なエンリッチメントデータを管理するための効率的で直感的なアプローチを提供します。
- SummArIzeRは、GitHubでユーザーマニュアルを備えたオープンソースRパッケージとして利用できます。
関連する概念動画
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K
Genome Annotation and Assembly
21.2K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.2K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
RNA-seq
12.3K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.3K
Amplifying Signals via Enzymatic Cascade
18.7K
When a ligand binds to a cell-surface receptor, the receptor's intracellular domain changes shape, which may either activate its enzyme function or allow its binding to other molecules. The initial signal is amplified by most signal transduction pathways. This means that a single ligand molecule can activate multiple molecules of a downstream target. Proteins that relay a signal are most commonly phosphorylated at one or more sites, activating or inactivating the protein. Kinases catalyze...
18.7K

