Related Experiment Video
Updated: May 5, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
Next-generation sequencing for molecular ecology: a caveat regarding pooled samples.
Eric C Anderson1, Hans J Skaug, Daniel J Barshis
1Fisheries Ecology Division, Southwest Fisheries Science Center, National Marine Fisheries Service, NOAA, 110 Shaffer Road, Santa Cruz, CA, 95060, USA; Department of Applied Math and Statistics (SOE2), University of California, 1156 High Street, Santa Cruz, CA, 95064, USA.
A new model using Dirichlet-compound multinomial distribution (CMD) predicts fixed single nucleotide polymorphism (SNP) loci. Analysis of pooled next-generation sequencing (NGS) data suggests limited gene copy representation, not population structure, causes excess fixed SNPs.
Area of Science:
- Population genetics
- Bioinformatics
- Genomics
Background:
- Next-generation sequencing (NGS) enables large-scale genetic analysis.
- Pooled sequencing reduces costs but can complicate data interpretation.
- Identifying population structure relies on accurately detecting genetic variation.
Purpose of the Study:
- To develop a model predicting fixed single nucleotide polymorphism (SNP) loci in pooled samples.
- To assess the contribution of population structure versus sampling artifacts to observed SNP patterns.
- To evaluate the reliability of using NGS on pooled samples for population genetics.
Main Methods:
- Development of a model based on Dirichlet-compound multinomial distribution (CMD) and Ewens sampling formula.
- Application of the model to analyze NGS data from Baltic Sea herring.
- Utilizing coalescent simulations to test hypotheses about population structure and sampling effects.
Main Results:
- The model predicted a fraction of fixed SNP loci between two pooled samples.
- Observed fixed loci in Baltic Sea herring data exceeded expectations under simple population structure.
- Coalescent simulations indicated that limited gene copy representation in pooled samples, rather than high population structure, likely explains the surplus of fixed SNPs.
Conclusions:
- The use of NGS on pooled samples for identifying divergent SNPs requires caution.
- Low read depth and limited individuals in pooled NGS data can lead to misinterpretations of population differentiation.
- Individually barcoded samples offer more reliable and potentially cost-effective data analysis for certain population genetics scenarios.
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...

