Exact word matches in rice pseudomolecules
Shaolin Liu1, Nicholas A Tinker, Diane E Mather
1Department of Plant Science, McGill University, Ste-Anne-de-Bellevue, Canada.
Abstract:
Using pseudomolecules of assembled genomic sequence, we computed the frequencies of 6 to 24 bp oligonucleotide (oligo) "words" across the genome of rice (Oryza sativa L. subsp. japonica). All oligos of 10 or fewer basepairs were repeated at least 12 times in the genome. The percentage of unique (non-repeated) oligos ranged from 0.1% for 12 bp oligos to 76.0% for 24 bp oligos. For three 200 kb regions, we annotated each nucleotide position with the genome-wide frequency of the 18 bp oligo starting at that position. These frequencies formed landscapes consisting of high- and low-frequency zones. Low-frequency zones contained occasional high-frequency spikes; these may represent footprints of RIM2 transposon activity. BLASTn searches of high-frequency non-SSR (simple sequence repeat) 18 bp oligos returned few sequences from species other than rice. These results demonstrate that, in rice, words are not randomly used between different regions within the same genome, and indicate that words that are frequently repeated within the rice genome tend to be unique to rice.
Related Concept Videos
¹H NMR: Pople Notation
A proton...
Riboswitches
The aptamer has high specificity for a particular metabolite which allows riboswitches to specifically regulate...
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons
Isomerism
Polymer Classification: Crystallinity
Crystalline domains are the regions where polymer chains are aligned in an orderly manner and held together in proximity by intermolecular forces. For example, chains in the crystalline domains of polyethylene and nylon are bound together by van der Waals...
Chirality at Nitrogen, Phosphorus, and Sulfur
A consequence of chirality is the need for enantiomeric resolution. While this is theoretically possible for all...


