Related Experiment Video
Updated: Jan 25, 2026

A Concoction Pipeline for Generating Molecular Operational Taxonomic Units (MOTUs) Among Riparian and Aquatic Beetles
Published on: July 11, 2025
An empirical pipeline for choosing the optimal clustering threshold in RADseq studies
Evan McCartney-Melstad1, Müge Gidiş2, H Bradley Shaffer1
1Department of Ecology and Evolutionary Biology, La Kretz Center for California Conservation Science, and Institute of the Environment and Sustainability, University of California, Los Angeles, California.
Abstract:
Genomic data are increasingly used for high resolution population genetic studies including those at the forefront of biological conservation. A key methodological challenge is determining sequence similarity clustering thresholds for RADseq data when no reference genome is available. These thresholds define the maximum permitted divergence among allelic variants and the minimum divergence among putative paralogues and are central to downstream population genomic analyses. Here we develop a novel set of metrics to determine sequence similarity thresholds that maximize the correct separation of paralogous regions and minimize oversplitting naturally occurring allelic variation within loci. These metrics empirically identify the threshold value at which true alleles at opposite ends of several major axes of genetic variation begin to incorrectly separate into distinct clusters, allowing researchers to choose thresholds just below this value. We test our approach on a recently published data set for the protected foothill yellow-legged frog (Rana boylii). The metrics recover a consistent pattern of roughly 96% similarity as a threshold above which genetic divergence and data missingness become increasingly correlated. We provide scripts for assessing different clustering thresholds and discuss how this approach can be applied across a wide range of empirical data sets.
Related Concept Videos
Choosing Between z and t Distribution
Empirical Method to Interpret Standard Deviation
This rule is used widely in statistics to calculate the proportion of data values...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Vesicular Tubular Clusters
With the help of motor proteins such...
Optimal Foraging
Optimization Problems

