Comparing the performance of selected variant callers using synthetic data and genome segmentation

Xiaopeng Bian1, Bin Zhu2, Mingyi Wang2

  • 1Center for Biomedical Informatics and Information Technology, National Cancer Institute, Rockville, MD, 20850, USA. bianxi@mail.nih.gov.

BMC Bioinformatics
|November 21, 2018
PubMed
Summary

Synthetic data with known mutations offers an economical way to validate genomic variant callers. An ensemble approach improved accuracy, providing a viable alternative to costly manual expert reviews for precision cancer medicine.

Related Concept Videos

Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes02:16

Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes

The present-day mitochondrial and chloroplast genomes have retained some of the characteristics of their ancestral prokaryotes and also have acquired new attributes during their evolution within eukaryotic cells. Like prokaryotic genomes, mitochondrial and chloroplast genomes neither bind with histone-like proteins nor show complex packaging into chromosome-like structures, as observed in eukaryotes. Unlike mitotic cell divisions observed in eukaryotic cells, mitochondria and chloroplasts...
16.1K
Selected Data About Geographic Locations01:25

Selected Data About Geographic Locations

Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
275
Genomics02:02

Genomics

Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
40.7K
Performing a Simple Data Analysis using MS-Excel Function01:17

Performing a Simple Data Analysis using MS-Excel Function

Microsoft Excel offers a suite of functions and tools ideal for statistical analysis, making it accessible to students and researchers. This article outlines fundamental Excel functions pivotal for data analysis.
SUM: This function calculates the total sum of a range of values. It's the foundation for aggregating data, essential for determining overall trends and totals in datasets.
AVERAGE: It computes the mean value of a given set of numbers, providing a quick insight into the central...
1.0K
Histone Variants at the Centromere02:30

Histone Variants at the Centromere

Histone variants are the histone proteins with structural and sequence variations. These variants may be regarded as “mutant” forms that replace their canonical histone counterparts in the nucleosomes. Specific post-translational modifications on the histone variants enable further chromatin complexity and regulate tissue-specific gene expression. The most common histone variants are from histone H2A, H2B, and linker histone H1 families. However, several variants of histone H3...
5.1K
Genome Size and the Evolution of New Genes03:21

Genome Size and the Evolution of New Genes

While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
9.1K