Related Experiment Video
Updated: Jun 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A large language model framework for sample-free population synthesis
Michael Jones1, Richard Dawson1, Jon Mills1
1School of Engineering, Newcastle University, Newcastle, United Kingdom.
This study introduces a novel framework using large language models (LLMs) to create synthetic populations for agent-based models without needing microdata. The LLM-based approach generates accurate, household-structured populations from aggregate data, enhancing accessibility for data-constrained research.
Area of Science:
- Computational Social Science
- Demographic Modeling
- Artificial Intelligence Applications
Background:
- Agent-based models (ABMs) require realistic synthetic populations for accurate simulations in diverse fields.
- Existing population synthesis methods often depend on scarce, privacy-restricted, or coarse-scale census microdata.
- This limitation restricts the application of ABMs in data-scarce environments.
Purpose of the Study:
- To present a sample-free framework for generating complete, household-structured synthetic populations using large language models (LLMs).
- To enable the creation of detailed demographic representations directly from aggregate data, overcoming microdata limitations.
- To expand the utility of ABMs in research settings with constrained data availability.
Main Methods:
- A multi-step, LLM-agnostic framework involving objective definition, input preparation, LLM selection, and synthetic household generation.
- Population synthesis is achieved through iterative prompting, where the LLM generates households guided by discrepancies between synthetic and target distributions.
- The method leverages the LLM's pre-trained knowledge for plausible attribute combinations, ensuring statistical alignment and structural feasibility without model fine-tuning.
Main Results:
- Global evaluation across 109 countries demonstrated high accuracy in reproducing marginal distributions like gender (SRMSE: 0.003) and household size (SRMSE: 0.026).
- Complex attributes such as household composition (SRMSE: 0.062) and age (SRMSE: 0.128) were also accurately reproduced.
- Case studies in Newcastle upon Tyne (UK) and Dar es Salaam (Tanzania) validated the framework's effectiveness.
Conclusions:
- The developed framework successfully generates coherent, household-structured synthetic populations from aggregate data, eliminating the need for microdata.
- This approach significantly enhances the applicability of agent-based modeling in diverse research areas facing data limitations.
- The LLM-based method offers an accessible and low-data-requirement solution for constructing foundational demographic datasets for simulations.
Related Concept Videos
Distributions to Estimate Population Parameter
Population Growth
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
What is Population Genetics?