Related Experiment Video
Updated: Jun 6, 2026

A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Benchmark datasets and software for developing and testing methods for large-scale multiple sequence alignment and
C Randal Linder1, Rahul Suri, Kevin Liu
1Integrative Biology, University of Texas; University of Texas; The University of Texas and Microsoft Research New England, and The Department of Computer Science, University of Texas at Austin.
Abstract:
We have assembled a collection of web pages that contain benchmark datasets and software tools to enable the evaluation of the accuracy and scalability of computational methods for estimating evolutionary relationships. They provide a resource to the scientific community for development of new alignment and tree inference methods on very difficult datasets. The datasets are intended to help address three problems: multiple sequence alignment, phylogeny estimation given aligned sequences, and supertree estimation. Datasets from our work include empirical datasets with carefully curated alignments suitable for testing alignment and phylogenetic methods for large-scale systematics studies. Links to other empirical datasets, lacking curated alignments, are also provided. We also include simulated datasets with properties typical of large-scale systematics studies, including high rates of substitutions and indels, and we include the true alignment and tree for each simulated dataset. Finally, we provide links to software tools for generating simulated datasets, and for evaluating the accuracy of alignments and trees estimated on these datasets. We welcome contributions to the benchmark datasets from other researchers.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Microbial Phylogeny
Modern Molecular Taxonomy
Applications of Molecular Taxonomy
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...

