Related Experiment Video
Updated: May 27, 2025

Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
Boosting GPT models for genomics analysis: generating trusted genetic variant annotations and interpretations through
1Department of Genetics, Yale University School of Medicine, New Haven, CT 06511, United States.
Large language models (LLMs) were enhanced for genomics using retrieval-augmented generation (RAG) and fine-tuning. RAG significantly improved variant annotation accuracy and accessibility, outperforming fine-tuning for factual knowledge injection.
Area of Science:
- Genomics
- Artificial Intelligence
- Bioinformatics
Background:
- Large language models (LLMs) possess broad knowledge but lack specialized domain expertise, such as in genomics.
- Variant annotation data is critical for interpreting genetic variants and prioritizing disease-related findings from large-scale sequencing.
- Current LLMs require enhancement to effectively process and interpret complex genomic data.
Purpose of the Study:
- To improve the performance of LLMs in the genomics domain.
- To integrate variant annotation data into LLMs using retrieval-augmented generation (RAG) and fine-tuning techniques.
- To enhance the interpretation and prioritization of genetic variants.
Main Methods:
- Implemented retrieval-augmented generation (RAG) to integrate 190 million variant annotations into GPT-4o.
- Utilized fine-tuning techniques on GPT-4 with variant annotation data.
- Compared the effectiveness of RAG and fine-tuning for knowledge injection.
Main Results:
- RAG successfully integrated a large volume of accurate variant annotations, enabling users to query specific variants for interpretation.
- Fine-tuning improved performance in some annotation fields, but overall accuracy remained suboptimal compared to RAG.
- RAG demonstrated superior performance over fine-tuning in terms of data volume, accuracy, and cost-effectiveness for factual knowledge integration.
Conclusions:
- The study successfully enhanced LLM capabilities in genomics through RAG and fine-tuning.
- RAG is a more effective method than fine-tuning for injecting factual genomic knowledge into LLMs.
- This pioneering work paves the way for advanced AI systems in genomics for clinical diagnosis and research.
More Related Videos
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016