Related Experiment Video
Updated: May 29, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Steering veridical large language model analyses by correcting and enriching generated database queries: first steps
1Department of Developmental and Cell Biology, Center for Complex Biological Systems, University of California at Irvine, 4203 McGaugh Hall, Irvine, CA 92697, USA.
Large language models (LLMs) struggle with genomics data accuracy. We introduce NagGPT to connect LLMs with databases, improving bioinformatics analysis and factual correctness.
Area of Science:
- Genomics and Bioinformatics
- Artificial Intelligence in Scientific Research
Background:
- Large language models (LLMs) possess factual knowledge but often lack completeness and retrieval accuracy, especially in rapidly evolving scientific fields like genomics.
- Existing LLMs, such as ChatGPT, exhibit limitations as bioinformatics assistants, including poor data retrieval, hallucination, and incorrect sequence manipulation.
Purpose of the Study:
- To address the limitations of LLMs in scientific domains by developing a system that grounds LLM outputs in current, authoritative data.
- To enhance LLM capabilities for data analysis in genomics and bioinformatics through improved factuality and instruction following.
Main Methods:
- Introduction of NagGPT, a middleware tool designed to interface between LLMs and databases, managing queries and responses.
- Development of a companion OpenAI custom GPT, Genomics Fetcher-Analyzer, which directs ChatGPT to generate and execute Python code for bioinformatics tasks using data from multiple genomics databases.
- Implementation of strategies to mitigate challenges such as code generation issues, identifier confusion, and data hallucination.
Main Results:
- Demonstration of NagGPT's effectiveness in bridging LLM knowledge gaps and facilitating database API usage.
- Successful execution of bioinformatics tasks by ChatGPT, guided by Genomics Fetcher-Analyzer and powered by dynamically retrieved data.
- Partial mitigation of issues related to code-data interaction, identifier ambiguity, and LLM hallucination.
Conclusions:
- The proposed system, incorporating NagGPT and Genomics Fetcher-Analyzer, offers a viable approach to augment LLMs as specialized bioinformatics assistants.
- Findings suggest pathways to enhance the factual accuracy and instruction-following capabilities of unmodified LLMs in scientific contexts.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Detection of Gross Error: The Q Test
Constraints and Statical Determinacy
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Common Leveling Mistakes and Errors

