Related Experiment Video
Updated: Oct 9, 2026

In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
Benchmarking commercial large language models for gene-disease-phenotype extraction from full-text human genetics
Danqing Yin1,2, Matthew Ka Siu Leung3, Darren Wan Ho Pun3
1School of Biomedical Sciences Li Ka Shing Faculty of Medicine The University of Hong Kong Hong Kong China.
Abstract:
Manual curation of gene-disease-phenotype relationships from the human genetics literature is a persistent bottleneck for maintaining its bioinformatics databases. Whereas large language models (LLMs) offer a promising alternative, there is currently no systematic benchmark that evaluates whether state-of-the-art commercial LLMs can perform this task reliably on the full-text articles. To address this gap, we introduce a standardized benchmark comprising 406 full-text articles covering 180 congenital heart disease-associated genes, and a multi-dimensional evaluation framework that incorporates fuzzy matching to account for synonyms and partial matches. We benchmarked seven state-of-the-art LLMs, GPT-4o, Claude-Opus-4, DeepSeek-R1, Grok-4, Qwen-3.5, Gemini-2.5 (Pro), and GPT-5 on the extraction of structured gene, disease, and phenotype fields. The top-performing model, Grok-4, achieved 97.6% overall accuracy, whereas the lowest-performing model reached approximately 88%, still surpassing many prior benchmarks employing zero-shot or n-shot prompting in biomedical relation extraction (RE) tasks. Our results provide a rigorous characterization of current LLMs capabilities and limitations. This paper contains two components. First, we conducted a human genetics field benchmark study on LLMs against a curated database. Second we developed the evaluation framework for this task. The benchmark dataset, evaluation framework, and model benchmarking outputs are made available online to support future studies in a reproducible manner.
More Related Videos
Related Concept Videos
Genomics
Incomplete Dominance
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Pharmacogenomics: Identification of New Drug Targets
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...

