Related Experiment Video
Updated: Aug 10, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Persona-Driven Data Augmentation for Disease Name Recognition Across Rare and General Disease Corpora: Comparative
Jude Crener Junior Pierre1, Tomohiro Nishiyama1, Shaowen Peng1
1Nara Institute of Science and Technology, Ikoma, Nara, Japan.
JMIR Medical Informatics
|July 24, 2026
Summary
Persona-driven data augmentation using large language models significantly improved biomedical disease named entity recognition (NER) performance. Combining multiple persona-generated variants with gold-standard data yielded the best results, especially for low-resource datasets.
Area of Science:
- Natural Language Processing
- Biomedical Informatics
- Machine Learning
Background:
- Medical information extraction, including named entity recognition (NER), is crucial but hindered by the scarcity and high cost of expert-annotated data.
- Standard data augmentation techniques often fail to preserve critical entity-label alignment in sequence-labeling tasks.
- Large language models (LLMs) can generate text but may introduce factual inconsistencies if not carefully controlled.
Purpose of the Study:
- To investigate the efficacy of persona-driven, document-level data augmentation using LLMs for enhancing biomedical disease NER.
- To assess if LLM-generated rephrasings can improve NER performance while preserving annotated entities.
- To evaluate the impact of diverse personas on NER model performance.
Main Methods:
- A data augmentation framework was developed using multiple personas with varying expertise, tone, and style.
- Personas rephrased training documents using XML-tagged prompts to maintain entity spans.
- The framework was evaluated on RareDis and NCBI disease NER datasets, measuring semantic fidelity (BERTScore) and lexical diversity (BLEU-4).
- Biomedical BioBERT models were fine-tuned and evaluated across different augmentation settings, including gold-standard data only, synonym replacement, and various persona combinations.
- Performance was assessed using microaveraged entity-level precision, recall, and F1-score.
Main Results:
- Persona-driven data augmentation consistently improved NER performance compared to gold-standard data alone across both datasets.
- The most significant gains were achieved by combining multiple persona-generated variants with gold-standard data.
- In the RareDis dataset, a low-fidelity persona subset yielded the best results (F1-score 73.35).
- In the NCBI disease dataset, the all-personas setting performed best (F1-score 89.32).
- For low-resource scenarios, certain persona augmentation settings in NCBI disease surpassed the performance of models trained on 100% gold-standard data using only 60% of the data.
- Analysis revealed improvements across entity types, notably for symptoms, and reduced symptom-sign confusion.
Conclusions:
- Persona-driven data augmentation effectively enhances biomedical disease NER by introducing controlled linguistic variation while preserving annotated entities.
- Combining multiple persona-generated variants with gold-standard data offers a promising strategy, particularly for low-resource biomedical NER tasks.
- The effectiveness of this approach varies across datasets, highlighting the need for careful persona selection and combination.

