Related Experiment Video
Updated: Jul 9, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
TADA: taxonomy-aware dataset aggregator
Emil Hägglund1, Siv G E Andersson1, Lionel Guy2
1Molecular Evolution, Department of Cell and Molecular Biology, Science for Life Laboratory, Biomedical Centre, Uppsala University, SE-751 24 Uppsala, Sweden.
Summary:
The profusion of sequenced genomes across the bacterial and archeal domains offers unprecedented possibilities for phylogenetic and comparative genomic analyses. In general, phylogenetic reconstruction is improved by the use of more data. However, including all available data is (i) not computationally tractable, and (ii) prone to biases, as the abundance of genomes is very unequally distributed over the biological diversity. Thus, in most cases, subsampling taxa to build a phylogeny is necessary. Currently, though, there is no available software to perform that handily. Here we present TADA, a taxonomic-aware dataset selection workflow that allows sampling across user-defined portions of the prokaryotic diversity with variable granularity, while setting constraints on genome quality and balance between branches.
Availability And Implementation:
TADA is implemented as a snakemake workflow and is freely available at https://github.com/emilhaegglund/TADA.
More Related Videos
Related Concept Videos
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Types of Aggregate Grading
Well-graded aggregates include a complete range of necessary size fractions that fit together to create a dense matrix with minimal voids, represented by a smooth, continuous gradation curve. This type of grading ensures good...
Mass Analyzers: Common Types
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...

