Related Experiment Video
Updated: Mar 16, 2026

09:37
Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
10.6K
Gene name errors are widespread in the scientific literature
Mark Ziemann1, Yotam Eren1,2, Assam El-Osta3,4
1Baker IDI Heart & Diabetes Institute, The Alfred Medical Research and Education Precinct, Melbourne, Victoria, 3004, Australia.
Genome Biology
|August 25, 2016
Summary
Microsoft Excel incorrectly converts gene names to dates and numbers. This common issue affects about 20% of genomics papers using Excel for gene lists, leading to errors in scientific data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Microsoft Excel's default settings can automatically convert gene names into dates or numerical formats.
- This automatic conversion poses a significant risk for data integrity in biological research.
Purpose of the Study:
- To assess the frequency of erroneous gene name conversions in supplementary data from genomics publications.
- To highlight the impact of spreadsheet software settings on biological data accuracy.
Main Methods:
- A programmatic analysis was conducted on supplementary gene lists from leading genomics journals.
- The study specifically identified instances where gene names were converted to non-standard formats.
Main Results:
- Approximately one-fifth (20%) of the analyzed papers contained gene lists with incorrect conversions.
- Common errors included the conversion of gene symbols to dates and floating-point numbers.
Conclusions:
- The default behavior of Microsoft Excel presents a widespread problem in genomics research.
- Researchers and journals should implement stricter data submission guidelines to prevent gene name conversion errors.
Related Concept Videos
Genetic Lingo
116.8K
Overview
116.8K
What is Genetic Engineering?
81.0K
Overview
81.0K
Exon Recombination
4.3K
The evolution of new genes is critical for speciation. Exon recombination, also known as exon shuffling or domain shuffling, is an important means of new gene formation. It is observed across vertebrates, invertebrates, and in some plants such as potatoes and sunflowers. During exon recombination, exons from the same or different genes recombine and produce new exon-intron combinations, which might evolve into new genes.
Exon shuffling follows “splice frame rules.” Each exon...
Exon shuffling follows “splice frame rules.” Each exon...
4.3K
Genome Copying Errors
5.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
5.3K
Leaky Scanning
5.8K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.8K
Genomic DNA in Eukaryotes
53.7K
Eukaryotes have large genomes compared to prokaryotes. To fit their genomes into a cell, eukaryotic DNA is packaged extraordinarily tightly inside the nucleus. To achieve this, DNA is tightly wound around proteins called histones, which are packaged into nucleosomes that are joined by linker DNA and coil into chromatin fibers. Additional fibrous proteins further compact the chromatin, which is recognizable as chromosomes during certain phases of cell division.
53.7K

