Related Experiment Video
Updated: May 9, 2025

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
PickMe: Sample Selection for Species Tree Reconstruction using Coalescent Weighted Quartets
Joseph Rusinko1, Yu Cai1, Allison Crysler1
1Department of Mathematics and Computer Science, Hobart and William Smith Colleges, Geneva, NY 14456, USA.
Abstract:
After collecting large datasets for phylogenomics studies, researchers must decide which genes or samples to include when reconstructing a species tree. Incomplete or unreliable datasets make the empiricist's decision more difficult. Researchers rely on ad hoc strategies to maximize sampling while ensuring sufficient data for accurate inferences. An algorithm called PickMe formalizes the sample selection process, assuming that the samples evolved under the tree Multispecies Coalescent Model. We propose a Bayesian framework for selecting samples for species tree analysis. Given a collection of gene trees, we compute a posterior probability for each quartet, describing the likelihood that the species tree displays this topology. From this, we assign individual samples reliability scores computed as the average of a scaled version of the posterior probabilities. PickMe uses these weights to recommend which samples to include in a species tree analysis. Analysis of simulated data showed that including the samples suggested by PickMe produced species trees closer to the true species trees than both unfiltered datasets and datasets with ad hoc gene occupancy cut-offs applied. To further illustrate the efficacy of this tool, we apply PickMe to gene trees generated from target capture data from milkweeds. PickMe indicates that more samples could have reliably been included in a previous milkweed phylogenomic analysis than the researchers analyzed without access to a formal methodology for sample selection. Using simulated and empirical data, we also compare PickMe to existing sample selection methods. Inclusion of PickMe will enhance phylogenomics data analysis pipelines by providing a formal structure for sample selection.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Speciation Rates
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Gene Evolution - Fast or Slow?
In contrast, regions which code...

