Related Experiment Video
Updated: Apr 22, 2026

11:23
Purifying the Impure: Sequencing Metagenomes and Metatranscriptomes from Complex Animal-associated Samples
Published on: December 22, 2014
34.6K
Metatranscriptomes from diverse microbial communities: assessment of data reduction techniques for rigorous
Andrew Toseland1, Simon Moxon, Thomas Mock
1School of Environmental Sciences, University of East Anglia, Norwich Research Park, Norwich, Norfolk NR4 7TJ, UK. a.toseland@uea.ac.uk.
BMC Genomics
|October 17, 2014
Summary
Data reduction techniques for metatranscriptome data impact annotation accuracy. Including unassembled reads alongside assembled contigs improves functional annotation, especially in diverse microbial communities.
Area of Science:
- Microbial Ecology
- Bioinformatics
- Genomics
Background:
- Metatranscriptome data often contains redundant sequences from diverse microbial populations, necessitating data reduction before annotation.
- Current data reduction techniques, such as clustering and de novo assembly, may impact the accuracy of taxonomic and functional annotation.
- The effect of these techniques on metatranscriptome data annotation and downstream interpretation remains unclear.
Purpose of the Study:
- To investigate the impact of common data reduction techniques on metatranscriptome data annotation.
- To compare the accuracy of clustering and de novo assembly methods for metatranscriptome data.
- To develop a simulation approach for assessing these techniques across varying microbial diversity.
Main Methods:
- Simulated metatranscriptome datasets using 454 and Illumina sequencing technologies with controlled taxonomic diversity.
- Assessed two data reduction methods: clustering and de novo assembly.
- Evaluated annotation accuracy based on protein domain content.
Main Results:
- For Illumina data, a two-step approach (assembly followed by clustering of contigs and unassembled sequences) yielded the most accurate protein domain annotation.
- For 454 data, combining annotations from both contigs and unassembled reads provided the most accurate protein domain annotations.
- The chosen data reduction strategy significantly influenced the accuracy of functional annotation.
Conclusions:
- Assembly should be attempted for metatranscriptome data, even in highly diverse environments.
- Unassembled reads should be included in the final annotation process to improve accuracy.
- These recommendations aim to provide a more accurate reflection of microbial transcriptional activity.
Related Concept Videos
Methods to Assess Microbial Communities
54
Microbial communities, comprising bacteria, archaea, and eukaryotic microorganisms, inhabit diverse ecosystems and play crucial roles in environmental and biological processes. Their diversity is defined by three main parameters: species richness (the number of distinct species), species abundance (the relative quantity of each species), and species evenness (how uniformly individual species are distributed in various locations). These factors together shape the structure and ecological balance...
54
Genome Annotation and Assembly
16.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
16.5K
Modern Molecular Taxonomy
828
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
828

