Related Experiment Video
Updated: Feb 26, 2026

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
ASQ-PHI: An adversarial synthetic data benchmark for clinical de-identification and search utility
James Weatherhead1, George Golovko2, Peter McCaffrey3
1Graduate School of Biomedical Sciences, University of Texas Medical Branch (UTMB), 301 University Boulevard, Galveston, TX 77555, USA.
A new synthetic dataset, ASQ-PHI, offers crucial resources for de-identifying clinical queries. This benchmark aids in safely transitioning Protected Health Information (PHI) from HIPAA-compliant large language models (LLMs) to external tools.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Data Science
Background:
- Hospitals utilize HIPAA-compliant large language models (LLMs) for clinical tasks, allowing Protected Health Information (PHI) input.
- LLMs require external data for up-to-date clinical evidence, necessitating a secure method for handling PHI during this transition.
- Existing de-identification datasets are not optimized for the short, query-style prompts used in chat-based clinical LLM interfaces.
Purpose of the Study:
- To introduce ASQ-PHI, a novel, fully synthetic benchmark dataset for de-identifying clinical queries in the context of LLM safe handoffs.
- To provide a resource for evaluating the de-identification of short, clinician-style search queries.
- To facilitate the development of robust systems for transitioning PHI-containing queries to external, non-BAA-covered tools.
Main Methods:
- Generated 1051 synthetic, single-turn clinical search queries mimicking clinician prompts for HIPAA-compliant LLMs.
- Annotated queries with PHI elements and their corresponding HIPAA Safe Harbor categories using machine-parsable delimiters and JSON.
- Employed an adversarial few-shot prompting pipeline with Azure OpenAI GPT-4o for query generation, including PHI-positive and hard-negative examples.
Main Results:
- The ASQ-PHI dataset contains 1051 queries, with 79.2% being PHI-positive and 20.8% hard negatives.
- 2973 PHI elements are labeled across 13 HIPAA Safe Harbor identifier types, suitable for measuring PHI removal and over-redaction.
- Baseline metrics for a commercial PHI detection service are provided within the dataset's repository.
Conclusions:
- ASQ-PHI addresses the scarcity of suitable datasets for de-identifying clinical queries in LLM safe handoff scenarios.
- The dataset enables the evaluation of de-identification performance on short, query-based prompts, distinct from traditional EHR narrative benchmarks.
- ASQ-PHI is released under an MIT license, promoting further research and development in secure clinical AI applications.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020