Related Experiment Video
Updated: Jan 11, 2026

Automated, Quantitative Cognitive/Behavioral Screening of Mice: For Genetics, Pharmacology, Animal Cognition and Undergraduate Instruction
Published on: February 26, 2014
Using Synthetic Data in Communication Sciences and Disorders to Promote Computational Reproducibility and
James C Borders1, Austin Thompson2, Elaine Kearney3,4
1Department of Speech, Language & Hearing Sciences, Boston University, MA.
Synthetic data can maintain statistical properties for communication sciences and disorders research, enhancing reproducibility. Further exploration is needed for hierarchical data structures.
Area of Science:
- Communication Sciences and Disorders
- Data Science
- Scientific Reproducibility
Background:
- Data sharing is crucial for scientific reproducibility but is uncommon in Communication Sciences and Disorders (CSD) due to privacy concerns.
- Synthetic data offers a solution by creating artificial datasets that mimic original data's statistical properties without compromising privacy.
Purpose of the Study:
- To explore the feasibility and preliminary utility of synthetic data generation for promoting transparency and reproducibility in CSD research.
- To assess if synthetic data can preserve statistical properties and inferential relationships from original CSD datasets.
Main Methods:
- Ten open CSD datasets from the American Speech-Language-Hearing Association
- Big Nine
- domains were utilized.
- Synthetic data were generated using the synthpop R package.
- General utility was assessed using visual inspection and the standardized ratio of propensity mean squared error (S_pMSE); specific utility was evaluated by comparing inferential statistics (model fit, coefficients, p-values) between original and synthetic data.
Main Results:
- All generated synthetic datasets demonstrated strong general utility, preserving univariate and bivariate distributions.
- Six out of nine synthetic datasets successfully maintained inferential relationships, indicating strong specific utility.
- Three datasets with hierarchical structures showed low specific utility, suggesting limitations for complex data.
Conclusions:
- Synthetic data effectively preserves statistical properties and relationships in non-hierarchical CSD data, supporting its use for enhancing reproducibility.
- Further research is required to develop effective methods for generating synthetic data from hierarchical structures.
- Researchers should evaluate the utility of synthetic data for their specific use cases to ensure accurate preservation of results.
More Related Videos
Related Concept Videos
Statistical Software for Data Analysis and Clinical Trials
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Data Reporting and Recording
Reliability and Validity
Ethics in Research
Censoring Survival Data

