Related Experiment Videos
Development and Evaluation of Large Language Model-Assisted Semi-Automated Data Harmonization Pipeline for the
Hyelee Kim1,2,3,4,5, Shuang Liang6,3, Kathy Lanier6,3
1Department of Epidemiology and Biostatistics, University of California, San Francisco, 550 16th St., Floor 2, San Francisco, US.
Background:
The increasing availability of machine-readable research data has created a growing need for efficient large-scale data harmonization (DH). Although large language models (LLMs) show promise for reducing the time and labor required for DH, their effective integration into harmonization workflows remains an important methodological challenge.
Objective:
To develop and evaluate a semi-automated, human-in-the-loop (HITL) LLM-assisted DH workflow as a proof of concept across eight projects within a research consortium focused on disparities in multiple chronic conditions.
Methods:
We developed an LLM-assisted DH workflow that combined preprocessing, semantic mapping, variable response mapping, synthetic data generation, data transformation, and iterative researcher review. Using GPT-4o hosted in a secure academic environment, project-specific survey items were mapped to 114 Consortium Common Data Element (CDE) semantic groups. Semantic mapping accuracy was evaluated based on the Consortium's consensus harmonization results. Complete and incomplete semantic mappings were distinguished according to construct overlap, and incomplete mappings requiring context-dependent human judgement were excluded for mapped item pairs and CDE semantic groups with and without HITL review.
Results:
Following preprocessing, 884 survey items were included in the DH workflow, of which 812 were mapped to a mean of 79 CDE semantic groups per project. Mean semantic mapping accuracy for item pairs was 95.9% with HITL review and 86.7% without HITL review. At the CDE semantic-group level, the corresponding accuracies were 98.9% and 94.9%, respectively. Across projects, omission of valid semantic mappings occurred more frequently than incorrect semantic mappings, particularly for heterogeneous response structures and subjective constructs measured using different instruments. Pipeline components-including concept-based preprocessing, iterative prompt refinement, and targeted HITL review-improved mapping completeness and supported variable harmonization across heterogeneous datasets.
Conclusions:
This proof-of-concept study demonstrates that the effectiveness of LLM-assisted DH depends more on the design of a structured HITL workflow rather than on the LLM alone. Combining preprocessing, iterative verification, and targeted human oversight improved the efficiency and reliability of DH while preserving researcher judgment for complex mapping decisions.