Related Experiment Video
Updated: Jun 12, 2025

10:28
Reusable Single Cell for Iterative Epigenomic Analyses
Published on: February 11, 2022
1.3K
Potential Benefits and Challenges of Quantifying Pseudoreplication in Genomic Data with Entropy Statistics
1Northwest Fisheries Science Center, National Marine Fisheries Service, National Oceanic and Atmospheric Administration, 2725 Montlake Blvd. East, Seattle, WA 98112, USA.
Entropy (Basel, Switzerland)
|September 27, 2024
Summary
Entropy metrics, like total correlation (TC), can help assess pseudoreplication in genetic data. This study found TC increases with more loci, offering insights for evolutionary ecology research.
Area of Science:
- Evolutionary Ecology
- Population Genetics
- Bioinformatics
Background:
- Generating large genetic marker datasets for evolutionary ecology is now common.
- Non-independent assortment of loci on chromosomes causes pseudoreplication, reducing precision in genetic analyses like linkage disequilibrium (LD).
- Overlapping locus pairs in LD analyses introduce another form of pseudoreplication.
Purpose of the Study:
- To explore the utility of entropy metrics, specifically total correlation (TC), for quantifying pseudoreplication in linkage disequilibrium (LD) studies.
- To investigate the relationship between TC, the number of loci (L), and effective population size (N).
Main Methods:
- Simulations were conducted on a monoecious population with varying effective population sizes (N) and numbers of loci (L).
- The study isolated the effect of overlapping locus pairs by analyzing unlinked loci.
- Entropy measures, particularly TC, were used to quantify inter-locus relationships.
Main Results:
- Total correlation (TC) showed a strong positive correlation with the number of loci (L).
- Effective population size (N) and other predictors had less pronounced effects on TC.
- Entropy-based metrics show promise for assessing statistical information in complex genetic datasets.
Conclusions:
- Entropy metrics, especially TC, can effectively quantify pseudoreplication in genetic data, particularly in LD studies.
- The number of loci significantly influences TC, highlighting its utility for managing genetic information.
- Scalability remains a challenge for large datasets due to computational limitations.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.6K
Statistical Analysis: Overview
6.2K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.2K
Gene Duplication and Divergence
6.1K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.1K
Genome Copying Errors
4.2K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.2K
Hardy-Weinberg Principle
71.9K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
71.9K
Epistasis Analysis
4.9K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
4.9K

