Related Experiment Video
Updated: Dec 15, 2025

A Concoction Pipeline for Generating Molecular Operational Taxonomic Units (MOTUs) Among Riparian and Aquatic Beetles
Published on: July 11, 2025
Phylogenetic double placement of mixed samples.
Metin Balaban1, Siavash Mirarab2
1Bioinformatics and Systems Biology Department, University of California San Diego, San Diego, CA 92093, USA.
This study introduces a new computational method for identifying organisms in mixed DNA samples and placing them phylogenetically. The MIxed Sample Analysis tool (MISA) successfully identifies constituent organisms even when absent from reference genomes.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying organisms in mixed samples is crucial for biodiversity, food safety, and evolutionary studies.
- Genome skimming generates low-coverage sequencing data, necessitating new methods for analyzing mixed samples without full genome assembly.
Purpose of the Study:
- To address the problem of phylogenetic double-placement: identifying organisms in a mixed sample and their phylogenetic positions.
- To develop a computational tool for analyzing mixed samples, particularly in the context of genome skimming.
Main Methods:
- Developed a model relating distances between mixed and reference samples using Jaccard indices of k-mer sets.
- Formalized phylogenetic double-placement as a non-convex optimization problem solvable numerically.
- Introduced the MIxed Sample Analysis tool (MISA) to implement the developed model.
Main Results:
- The MISA tool successfully identifies constituent organisms in mixed samples.
- The method performs well on simulated and biological datasets, despite underlying assumptions.
- Phylogenetic placement is achieved simultaneously with mixture decomposition.
Conclusions:
- The developed model and MISA tool offer an effective solution for the phylogenetic double-placement problem.
- This approach has broad applications in biodiversity research, food provenance, and evolutionary reconstruction.
- The method demonstrates robust performance in practical scenarios using genome skimming data.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
Phylogenetic Trees
Phylogeny
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...

