Federated knowledge retrieval elevates large language model performance on biomedical benchmarks

Janet Joy1, Andrew I Su1

  • 1Department of Integrative Structural and Computational Biology, Scripps Research, 10550 N Torrey Pines Rd, La Jolla, CA, 92037, USA.

Gigascience
|January 19, 2026
PubMed
Summary

Retrieval-augmented generation using BioThings Explorer (BTE-RAG) enhances large language model (LLM) accuracy in biomedical research. This framework improves factual correctness and mechanistic exploration for drug discovery and translational science.

Related Concept Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

This article describes RUGGED (Retrieval Under Graph-Guided Explainable disease Distinction), which integrates Large Language Model (LLM) inference with Retrieval-Augmented Generation (RAG). It draws evidence from expert-curated biomedical knowledge bases and peer-reviewed biomedical publications to synthesize new knowledge from up-to-date information, identify explainable and actionable predictions, and pinpoint promising directions for hypothesis-driven...
1.3K
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

A computational protocol, CaseOLAP LIFT, and a use case are presented for investigating mitochondrial proteins and their associations with cardiovascular disease as described in biomedical reports. This protocol can be easily adapted to study user-selected cellular components and...
2.1K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

In this protocol, foundation large language model response quality is improved via augmentation with peer-reviewed, domain-specific scientific articles through a vector embedding mechanism. Additionally, code is provided to aid in performance comparison across large language...
1.0K
The (Spatial) Memory Game: Testing the Relationship Between Spatial Language, Object Knowledge, and Spatial Cognition05:15

The (Spatial) Memory Game: Testing the Relationship Between Spatial Language, Object Knowledge, and Spatial Cognition

We present a protocol to explore the relationship between spatial language production, spatial memory, and object knowledge. The procedure allows experimental manipulation of, and control over, conditions of object knowledge, language at instruction, and physical location, thus teasing apart cognitive and linguistic models describing interactions between these...
11.3K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Existing algorithms generate one solution for a biomarker detection dataset. This protocol demonstrates the existence of multiple similarly effective solutions and presents a user-friendly software to help biomedical researchers investigate their datasets for the proposed challenge. Computer scientists may also provide this feature in their biomarker detection...
8.0K
Retrieval01:12

Retrieval

Retrieval is the process of getting information out of memory storage and back into conscious awareness. This ability is essential for daily tasks like brushing hair and teeth, driving to work, and performing job duties. Retrieval occurs in three ways: recall, recognition, and relearning.
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
418