Related Experiment Video
Updated: Jun 5, 2025

IR-TEx: An Open Source Data Integration Tool for Big Data Transcriptomics Designed for the Malaria Vector Anopheles gambiae
Published on: January 15, 2020
Contamination Survey of Insect Genomic and Transcriptomic Data.
Jiali Zhou1, Xinrui Zhang1, Yujie Wang1
1State Key Laboratory of Ecological Pest Control for Fujian and Taiwan Crops, College of Plant Protection, Fujian Agriculture and Forestry University, Fuzhou 350002, China.
This study investigated contamination in insect genomic and transcriptomic data, finding higher rates in transcriptomic assemblies (TSA) than whole-genome sequencing (WGS) data. Researchers identified contamination causes and proposed mitigation strategies for public databases.
Area of Science:
- Genomics
- Bioinformatics
- Zoology
Background:
- High-throughput sequencing generates vast amounts of data, increasing the risk of contamination from non-target species.
- Insecta, a highly diverse arthropod group, requires thorough evaluation of contamination in public genomic and transcriptomic databases.
- Understanding contamination sources is crucial for accurate biodiversity and evolutionary studies.
Purpose of the Study:
- To evaluate contamination prevalence in insect whole-genome sequencing (WGS) and transcriptomic (TSA) data from GenBank.
- To analyze potential causes of insect and mammal contamination in genomic datasets.
- To propose a workflow for detecting and mitigating contamination in sequencing data.
Main Methods:
- Utilized Cytochrome c Oxidase subunit I (COI) barcodes for contamination detection.
- Analyzed 2796 WGS and 1382 TSA assemblies across four major insect orders.
- Investigated contamination from insects and mammals within the datasets.
Main Results:
- Contamination was detected in 1.14% of WGS and 11.0% of TSA assemblies.
- TSA data showed significantly higher contamination rates compared to WGS data.
- Contamination varied by insect order: Hemiptera (9.22%), Hymenoptera (7.66%), Coleoptera (3.48%), and Diptera (1.89%).
- Potential contamination sources including food, parasitism, and cross-contamination were identified.
Conclusions:
- Transcriptomic data present a higher contamination risk than whole-genome data in insects.
- Contamination levels are order-specific, necessitating tailored detection methods.
- A standardized workflow and mitigation strategies are proposed to improve data quality in public repositories.
More Related Videos
08:36Empirical, Metagenomic, and Computational Techniques Illuminate the Mechanisms by which Fungicides Compromise Bee Health
Published on: October 9, 2017
07:20Maintaining Biological Cultures and Measuring Gene Expression in Aphis nerii: A Non-model System for Plant-insect Interactions
Published on: August 31, 2018