Related Experiment Video
Updated: Jul 16, 2025

09:35
A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
17.9K
Evaluation of large language models for discovery of gene set function
Mengzhou Hu1, Sahar Alkhairy2, Ingoo Lee1
1Department of Medicine, University of California San Diego, La Jolla, California, USA.
Arxiv
|September 21, 2023
Summary
Large Language Models (LLMs) can assist in functional genomics by identifying gene functions from
Area of Science:
- Functional genomics
- Bioinformatics
- Computational biology
Background:
- Current gene function databases are incomplete, limiting functional genomics.
- Gene set analysis requires accurate identification of common biological functions.
- Evaluating the utility of Large Language Models (LLMs) for gene set analysis is crucial.
Approach:
- Assessed five LLMs (GPT-4, Gemini-Pro, Mixtral-Instruct, Llama2-70b) for gene function discovery.
- Benchmarked LLM performance against curated Gene Ontology gene sets and random gene sets.
- Evaluated LLM ability to provide rationale, citations, and confidence scores for identified functions.
- Tested LLMs on gene sets derived from 'omics data to identify novel functions.
Key Points:
- GPT-4 demonstrated high confidence in identifying functions for canonical gene sets (73%) and zero confidence for random sets.
- Gemini-Pro and Mixtral-Instruct showed naming ability but were overconfident on random sets; Llama2-70b performed poorly.
- GPT-4 identified novel, verifiable gene functions from 'omics data not found by traditional methods (32% of cases).
Conclusions:
- LLMs, particularly GPT-4, show significant promise as tools for functional genomics and 'omics data analysis.
- LLMs can rapidly synthesize common gene functions, acting as valuable assistants for researchers.
- Further development and validation are needed, but LLMs offer a powerful new approach to supplement existing functional enrichment methods.
Related Concept Videos
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Genome Size and the Evolution of New Genes
2.5K
2.5K
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K

