Related Experiment Video
Updated: Apr 8, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Data biases in genomics
Lusine Nazaretyan1, Martin Kircher2
1Exploratory Diagnostic Sciences, Berlin Institute of Health at Charité - Universitätsmedizin Berlin, Berlin 10117, Germany.
None:
Machine learning (ML) is developing into an inherent part of genomic research due to the ever-increasing amounts of genomic data. However, data-driven algorithms are strongly dependent on good quality and representative data, which can be problematic in genomics due to various reasons. One of these reasons is data biases-flawed or incomplete data often containing systematic errors that compromise its representativeness. In this review, we examine different categories of data biases in genomics and translate them into the framework of general ML. We give examples of different types of biases present in widely used databases such as NCBI ClinVar and gnomAD and illustrate how data biases can influence model performance in assorted studies.
Related Concept Videos
Genomics
Genomic Imprinting and Inheritance
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
Incomplete Dominance
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Bias in Epidemiological Studies
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...

