Related Experiment Video
Updated: Jun 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Scaling sensor metadata extraction for exposure health using LLMs
Fatemeh Shah-Mohammadi1, Sunho Im2, Julio C Facelli1,3
1Department of Biomedical Informatics, The University of Utah, Salt Lake City, UT 84108, United State.
Background:
The rapid evolution and diversity of sensor technologies, coupled with inconsistencies in how sensor metadata is reported across formats and sources, present significant challenges for generating exposomes and exposure health research.
Objective:
Despite the development of standardized metadata schemas, the process of extracting sensor metadata from unstructured sources remains largely manual and unscalable. To address this bottleneck, we developed and evaluated a large language model (LLM)-based pipeline for automating sensor metadata extraction and harmonization from publicly available exposure health literature.
Methods:
Using GPT-4 in a zero-shot setting, we constructed a pipeline that parses full-text PDFs to extract metadata and harmonizes output into structured formats.
Results:
Our automated pipeline achieved substantial efficiency gains in completing extractions much faster than manual review and demonstrated strong performance with 88.0% accuracy, 88.0% precision, 93.0% recall, and an F1-score of 90.0%.
Conclusions:
This study demonstrates the feasibility and scalability of leveraging LLMs to automate sensor metadata extraction for exposure health, reducing manual burden while enhancing metadata completeness and consistency. Our findings support the integration of LLM-driven pipelines into exposure health informatics platforms.