Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Pre-Meta: priors-augmented retrieval for LLM-based metadata generation
Phil Tinn1, Sondre Sørbø1, Shanshan Jiang1
1SINTEF AS, Oslo 0373, Norway.
Pre-Meta enhances automated metadata generation for genomic datasets using LLMs. This pipeline improves annotation accuracy, facilitating better data discovery and publication across repositories.
Area of Science:
- Genomics
- Bioinformatics
- Data Science
Background:
- High-throughput sequencing generates vast genomic data, but manual annotation and metadata creation hinder discovery and publication.
- Large language models (LLMs) show promise for streamlining dataset profiling, yet struggle with specialized domains like biomedical genomics.
- Current limitations in LLM generalization impede efficient use of genomic data resources.
Purpose of the Study:
- To present Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline.
- To improve automated metadata generation accuracy by leveraging related priors like metadata tags and ontologies.
- To enhance the discovery and publication of genomic data resources.
Main Methods:
- Developed Pre-Meta, a data annotation pipeline.
- Implemented an enriched retrieval procedure using auxiliary information (metadata tags, ontologies).
- Validated the pipeline on five metadata fields across 1500 papers.
Main Results:
- Pre-Meta demonstrated systemic improvement in the annotation task without finetuning or prompt optimization.
- Achieved accuracy gains of 23% (GPT-4o mini), 72% (Llama 8B), and 75% (Mistral 7B) over conventional RAG.
- Validated the effectiveness of leveraging related priors for metadata generation.
Conclusions:
- Pre-Meta significantly enhances automated metadata generation for genomic datasets.
- The pipeline improves LLM performance in specialized domains by utilizing prior knowledge.
- Facilitates more efficient discovery and publication of genomic data resources.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Sources of Law
Constitutional law is foundational, deriving from federal and state constitutions, and...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Methods of Documentation II: POMR
Archival Research
Law of Independent Assortment