Related Experiment Video
Updated: Jan 13, 2026

Semiconductor Sequencing for Preimplantation Genetic Testing for Aneuploidy
Published on: August 25, 2019
Synthetic data-driven AI approach for fetal chromosomal aneuploidies detection
Changhoe Hwang1, Krishna Prasad Adhikari1, Gyeongin Oh1
1Department of Research & Development, Genomecare Inc., Suwon, Gyeonggi, 16229, Korea.
Motivation:
A major limitation in the development of fetal chromosomal aneuploidy detection technologies lies in the scarcity of real positive data. To address this issue, we propose a novel methodology to generate virtually unlimited synthetic negative and positive datasets with >99.9% similarity to real data, enabling accurate detection of both autosomal chromosome aneuploidies (ACA) and sex chromosome aneuploidies (SCA). In terms of methods, blood samples from 15 999 pregnant women were analyzed, including 186 clinically confirmed positive cases. Using 701 high-confidence negatives as a reference, we designed algorithms for synthetic data generation. For negatives, multiple real FASTQ files were randomly merged, and fetal fraction (FF) was recalculated to reflect biological variability. For positives, chromosome-specific read counts were adjusted using numerical equations: ACAs were simulated by increasing the target chromosome reads, and SCAs were generated by adjusting sex chromosome read counts using regression models that account for FF and total read count, with the GC distribution preserved. Logistic regression (LR) models were then trained using features including FF, GC content, and chromosomal read counts. Performance was evaluated against conventional z-score methods and real positive cases.
Results:
From high-confidence negative samples, ∼160 000 synthetic training datasets were generated for major ACA and ∼35 000 for each SCA. While z-score methods showed declines in sensitivity (T13) or positive predictive value (PPV) (T18, T21) under low prevalence, LR models consistently maintained 100% sensitivity and PPV for ACAs, achieved ≥99.6% sensitivity and PPV for SCAs on synthetic evaluation datasets, and demonstrated 100% accuracy on real positive samples.
Availability And Implementation:
All formulas and procedures required for synthetic data generation and model development are implemented in Python and are available at https://github.com/genomecare-rnd/SyntheticData-NIPT.
Related Concept Videos
Nondisjunction
Karyotyping

