Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Modeling and Similitude01:12

Modeling and Similitude

246
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
246
Selected Data About Geographic Locations01:25

Selected Data About Geographic Locations

26
Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
26
Stratified Sampling Method01:16

Stratified Sampling Method

11.7K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
11.7K
GIS Software, Hardware, and Sources of GIS Data01:23

GIS Software, Hardware, and Sources of GIS Data

46
A Geographic Information System (GIS) combines specialized software and hardware to effectively manage, analyze, and present spatial and related data. GIS software includes critical functionalities such as a user interface for easy navigation, database management tools for handling spatial and attribute data, and data retrieval features for efficient access. Analytical tools transform raw data into insights, while display functions produce maps and reports in various formats for effective...
46
Cluster Sampling Method01:20

Cluster Sampling Method

11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Levels of Use of a GIS01:29

Levels of Use of a GIS

45
Geographic Information Systems (GIS) operate across three levels of application, each representing an increasing degree of complexity: data management, analysis, and prediction. These levels reflect the expanding functionality and versatility of GIS technology in handling spatial data for diverse purposes.Data ManagementAt its foundational level, GIS serves as a tool for data management, enabling the input, storage, retrieval, and organization of spatial data. This level is often employed in...
45

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Letter of Concern About Ruuska et al. From 4 April 2026.

Acta paediatrica (Oslo, Norway : 1992)·2026
Same author

Clinical trial simulation of antiviral drugs.

Journal of virology·2026
Same author

Prevalence and Health Care Disparities of Retinal Conditions: A Meta-Analysis.

JAMA ophthalmology·2026
Same author

Towards an estimate of the impact of censorship on biomedical literature.

Journal of the American Medical Informatics Association : JAMIA·2025
Same author

Comparing 3 Evidence-Based Strategies to Reduce Cardiovascular Disease Burden: An Individual-Based Cardiometabolic Policy Simulation.

Journal of the American Heart Association·2025
Same author

Optimal allocation of antenatal and young child nutrition interventions: an individual-based global burden of disease calibrated microsimulation.

BMC global and public health·2025

Related Experiment Video

Updated: Jun 9, 2025

A Highly Scalable Approach to Perform Ecological Surveys of Selfing Caenorhabditis Nematodes
09:10

A Highly Scalable Approach to Perform Ecological Surveys of Selfing Caenorhabditis Nematodes

Published on: March 1, 2022

2.5K

Simulated data for census-scale entity resolution research without privacy restrictions: a large-scale dataset

Beatrix Haddock1, Alix Pletcher1, Nathaniel Blair-Stahn1

  • 1Institute for Health Metrics and Evaluation, University of Washington, Seattle, Washington, 98195, USA.

Gates Open Research
|October 30, 2024
PubMed
Summary

The pseudopeople package generates realistic, simulated population data for entity resolution (ER) research. This enables algorithm development without using sensitive personal information, overcoming data access barriers in data science.

Keywords:
Entity resolution (ER)microsimulation

More Related Videos

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.4K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

8.7K

Related Experiment Videos

Last Updated: Jun 9, 2025

A Highly Scalable Approach to Perform Ecological Surveys of Selfing Caenorhabditis Nematodes
09:10

A Highly Scalable Approach to Perform Ecological Surveys of Selfing Caenorhabditis Nematodes

Published on: March 1, 2022

2.5K
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.4K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

8.7K

Area of Science:

  • Data Science
  • Computational Social Science

Background:

  • Entity Resolution (ER) is crucial for data integration but hindered by privacy concerns surrounding personally identifiable information (PII).
  • Restrictions on accessing authentic PII slow the development and testing of new ER methods and software.
  • The pseudopeople Python package addresses this by generating simulated datasets for ER research.

Purpose of the Study:

  • To develop and release a tool for generating realistic, noisy simulated population data for entity resolution (ER).
  • To enable researchers to develop and test ER algorithms without compromising sensitive personal information.

Main Methods:

  • Utilized the Vivarium simulation platform to create a dynamic model of individuals, families, households, and employment.
  • Generated simulated censuses, surveys, and administrative data reflecting real-world population dynamics.
  • Developed the pseudopeople Python package to add configurable noise to simulated data for realistic ER challenges.

Main Results:

  • Produced over 900 gigabytes of simulated population data, representing hundreds of millions of individuals.
  • Made a sample population of thousands openly available via the pseudopeople package.
  • Provided access to larger simulated populations upon request, structured for ER research.

Conclusions:

  • The pseudopeople package and its associated simulated data overcome PII access barriers in ER research.
  • Facilitates the development, testing, and adoption of novel ER algorithms and software for large-scale population data.