Related Experiment Video
Updated: May 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility
Hanshu Rao1, Weisi Liu1, Haohan Wang2
1Department of Computer Science, University of Memphis, Memphis, 38152 TN United States.
Journal of Healthcare Informatics Research
|May 11, 2026
Summary
Synthetic data generation using large language models (LLMs) shows promise for biomedical research. LLMs address data scarcity and quality, but standardized evaluation and accessibility remain key challenges for future development.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Healthcare
- Data Science
Background:
- Biomedical research faces challenges with data scarcity, utility, and quality.
- Large Language Models (LLMs) are increasingly adopted for synthetic data generation in this field.
- Existing reviews lack a systematic focus on LLM-driven synthetic data for biomedical applications.
Purpose of the Study:
- To systematically review recent advances in LLM-based synthetic data generation for biomedical and clinical research.
- To analyze how LLMs address data scarcity, utility, and quality across different data modalities.
- To identify current limitations and future directions in this rapidly evolving area.
Main Methods:
- Conducted a scoping review following PRISMA-ScR guidelines.
- Searched major academic databases (PubMed, ACM, Web of Science, Google Scholar) for literature from 2020-2025.
- Included 59 relevant studies focusing on synthetic data generation in biomedical contexts.
Main Results:
- Unstructured text was the predominant data modality (78.0%), followed by tabular (13.6%) and multimodal data (8.4%).
- LLM prompting (74.6%) was the most common generation method, followed by fine-tuning (20.3%).
- Evaluation methods were diverse, with human-in-the-loop assessments (44.1%) being most frequent, followed by intrinsic metrics (27.1%).
Conclusions:
- LLMs offer significant potential for generating synthetic biomedical data, particularly for unstructured text.
- Key barriers include data modality limitations, domain-specific utility, resource accessibility, and the need for standardized evaluation protocols.
- Future research should prioritize developing transparent evaluation frameworks and enhancing accessibility to facilitate widespread adoption in biomedical research.
More Related Videos
Related Concept Videos
Synthetic Biology
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...

