Related Experiment Video
Updated: Nov 12, 2025

09:03
Profiling Individual Human Embryonic Stem Cells by Quantitative RT-PCR
Published on: May 29, 2014
11.8K
Tracing foreign sequences in plant transcriptomes and genomes using OCT4, a POU domain protein
Adeleh Saffar1, Maryam M Matin2,3
1Department of Biology, Faculty of Science, Ferdowsi University of Mashhad, Mashhad, Iran.
Molecular Genetics and Genomics : MGG
|March 19, 2021
Summary
Sequencing data, including plant transcriptomes and reference genomes, can be contaminated with foreign RNA, such as arthropod POU domain proteins and rDNA. This contamination can lead to errors in downstream analyses and misrepresentation of molecular interactions.
Area of Science:
- Genomics and Bioinformatics
- Molecular Biology
- Computational Biology
Background:
- Contaminants in sequencing data, particularly reference genomes and transcriptomes, introduce errors into downstream analyses.
- Such contaminants can misrepresent the molecular basis of biological interactions.
- POU domain proteins are a family of proteins not previously reported in plants and fungi.
Purpose of the Study:
- To report the presence of plant transcriptomes contaminated with RNAs encoding POU domain proteins.
- To identify foreign DNA sequences within plant reference genomes and draft genomes.
- To highlight the implications of such contaminations for genomic and transcriptomic research.
Main Methods:
- Bioinformatic analysis of plant transcriptome and genome sequencing data.
- Identification and characterization of foreign RNA and DNA sequences.
- Comparative genomics to trace the origin of contaminant sequences.
Main Results:
- A significant number of plant transcriptomes were found to be contaminated with RNAs encoding POU domain proteins.
- The reference genome of Rhodamnia argentea contained four POU domain protein-coding sequences of arthropod origin.
- Draft genomes of Humulus lupulus and Cannabis sativa contained complete rDNA sequences from Tetranychus species (arthropods).
Conclusions:
- Foreign fragments, including arthropod-related sequences, are present in plant sequencing data.
- These contaminants can be misidentified as endogenous elements, impacting biological interpretation.
- Thorough screening of sequencing data for contaminants before public release and checking existing genomes is crucial.
More Related Videos
Related Concept Videos
Transgenic Plants
8.0K
Recombinant DNA technology called transgenesis is often used to add a foreign gene or remove a detrimental gene from an organism. Such genetically modified organisms are called transgenic organisms.
The first-ever transgenic plant was a tobacco plant developed in 1983 that showed resistance against the tobacco mosaic virus. Since then, many transgenic plants have been developed and commercialized for improving the agricultural, ornamental, and horticultural value of a crop plant. Transgenic...
The first-ever transgenic plant was a tobacco plant developed in 1983 that showed resistance against the tobacco mosaic virus. Since then, many transgenic plants have been developed and commercialized for improving the agricultural, ornamental, and horticultural value of a crop plant. Transgenic...
8.0K
Transposons
551
Transposons, or "jumping genes," are small mobile genetic elements (MGEs) that range from 700 to 40,000 base pairs in length. They are found in all organisms and can move within the same chromosome or transfer to different chromosomes. In some cases, transposons can also jump between different host DNA molecules, such as plasmids or viruses, contributing to genetic variability.Barbara McClintock first discovered these mobile genetic elements in the 1940s while studying maize genetics, and she...
551
DNA-only Transposons
15.5K
DNA-only transposons are called autonomous transposons since they code for the enzyme transposase that is required for the transposition mechanism. Insertion of transposons can alter gene functions in multiple ways. They can mutate the gene, alter gene expression by introducing a novel promoter or insulator sequence, introduce new splice sites, and change the mRNA transcripts produced, or remodel chromatin structure.
The donor site from where the transposon is excised is either degraded or...
The donor site from where the transposon is excised is either degraded or...
15.5K
Genomic DNA in Eukaryotes
50.7K
Eukaryotes have large genomes compared to prokaryotes. To fit their genomes into a cell, eukaryotic DNA is packaged extraordinarily tightly inside the nucleus. To achieve this, DNA is tightly wound around proteins called histones, which are packaged into nucleosomes that are joined by linker DNA and coil into chromatin fibers. Additional fibrous proteins further compact the chromatin, which is recognizable as chromosomes during certain phases of cell division.
50.7K
DNA Microarrays
19.5K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
19.5K
Combinatorial Gene Control
8.9K
Combinatorial gene control is the synergistic action of several transcriptional factors to regulate the expression of a single gene. The absence of one or more of these factors may lead to a significant difference in the level of gene expression or repression.
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
8.9K

