Related Experiment Video
Updated: Jan 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automating candidate gene prioritization with large language models: from naive scoring to literature-grounded
Taushif Khan1, Mohammed Toufiq1, Marina Yurieva1
1The Jackson Laboratory for Genomic Medicine, 10 Discovery Drive, Farmington, Connecticut, 06032, United States.
This study presents a novel framework for prioritizing therapeutic gene targets using large language models (LLMs) and literature validation. The approach successfully identifies sepsis-relevant genes and novel therapeutic candidates, overcoming LLM limitations for reliable biological insights.
Area of Science:
- Biomedical research
- Computational biology
- Genomics
Background:
- Transcriptomic studies generate vast gene data, posing challenges for identifying therapeutic targets.
- Large language models (LLMs) show promise for gene prioritization but often hallucinate and lack validation.
- Systematic validation against expert knowledge is crucial for reliable gene prioritization.
Purpose of the Study:
- To develop and validate a computational framework for systematic gene prioritization using LLMs and literature.
- To identify high-confidence therapeutic gene targets for sepsis.
- To overcome the limitations of LLMs in biomedical research through rigorous validation.
Main Methods:
- A two-stage framework combining LLM-based screening with literature validation was developed.
- Genes were screened for sepsis relevance using multi-criteria evaluation and retrieval-augmented generation from sepsis publications.
- A faithfulness evaluation system ensured LLM predictions aligned with literature evidence.
Main Results:
- The framework identified 609 sepsis-relevant genes with high filtering efficiency, enriched in inflammatory pathways.
- 30 ultra-high confidence therapeutic candidates were identified, including novel targets and known sepsis genes.
- Benchmark validation achieved 71.2% recall, correlating computational confidence with evidence quality.
Conclusions:
- The framework transforms unreliable LLM outputs into systematically validated biological insights.
- It provides a practical tool for prioritizing experimental validation and literature-guided biomarker discovery.
- The modular design allows adaptation to other diseases, offering a versatile approach.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Quantifying and Rejecting Outliers: The Grubbs Test
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

