Related Experiment Video
Updated: May 8, 2026

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved (Non-model) Organisms
Published on: May 9, 2017
All Models are Wrong, Some are Annotated: Automating Metadata in Biomedical Repositories
Inessa Cohen1, Hongyi Yu2, Robert A McDougal1,2,3,4,5
1Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT 06510, United States.
Large language models (LLMs) can automatically generate metadata for scientific models from source code, improving discovery. These AI tools show promise for annotating large biomedical repositories more efficiently than traditional methods.
Area of Science:
- Computational Neuroscience
- Bioinformatics
- Artificial Intelligence
Background:
- High-quality metadata is crucial for scientific discovery but is often sparse in large data repositories.
- Manually annotating complex biological details, such as ion channel and receptor subtypes, from source code is time-consuming and challenging.
Purpose of the Study:
- To evaluate the effectiveness of large language models (LLMs) in automatically inferring ion channel and receptor subtype metadata directly from source code.
- To compare LLM performance against a traditional feature-engineered baseline model.
Main Methods:
- Extracted 5,133 model files from the ModelDB repository.
- Manually annotated 1,100 models, with 253 reserved for testing.
- Evaluated LLM approaches (GPT-5.2, GPT-mini) using zero-shot and heuristic-augmented prompting.
- Compared LLM performance (accuracy, precision, recall, F1 score) against an XGBoost baseline model.
Main Results:
- LLMs significantly outperformed the XGBoost baseline model in metadata annotation.
- Heuristically augmented GPT-mini achieved 96.0% accuracy at the type level and 88.1% accuracy at the subtype level.
- LLM outputs were consistent and errors were generally limited to related biological families.
Conclusions:
- LLMs show strong potential for scalable metadata generation directly from scientific source code, requiring minimal tuning.
- While effective, LLM performance can vary across subtypes, necessitating domain-specific validation and careful evaluation.
- The approach is promising for enhancing biomedical repositories and may generalize to other scientific code repositories.
More Related Videos
Related Concept Videos
Genome Annotation and Assembly
Mechanistic Models: Compartment Models in Individual and Population Analysis
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Mismatch Repair
Molecular Models
Mechanistic Models: Overview of Compartment Models

