Related Experiment Videos
Mistaken identifiers: gene name errors can be introduced inadvertently when using Excel in bioinformatics
Barry R Zeeberg1, Joseph Riss, David W Kane
1Genomics & Bioinformatics Group, Laboratory of Molecular Pharmacology, Center for Cancer Research (CCR), National Cancer Institute (NCI), National Institutes of Health (NIH), Bethesda, MD 20892 USA. barry@discover.nci.nih.gov
BMC Bioinformatics
|June 25, 2004
Summary
Excel software inadvertently alters gene names to non-gene names during data processing. This issue affects thousands of genes, including medically important ones, and requires user awareness and workarounds.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Microarray data analysis can be compromised by unintended gene name alterations.
- Gene names are sometimes converted to non-gene names during data processing.
Purpose of the Study:
- To identify the cause of inadvertent gene name conversions in data processing.
- To inform users about potential data integrity issues in gene name analysis.
Main Methods:
- Investigated data processing anomalies in microarray datasets.
- Traced gene name conversion errors to default Microsoft Excel formatting settings.
Main Results:
- Default date and floating-point format conversions in Excel alter gene names.
- At least 30 gene names are affected by date conversions and over 2,000 by floating-point conversions.
- These irreversible conversions can lead to the loss of gene information.
Conclusions:
- Excel users must be aware of automatic formatting that corrupts gene names.
- This issue impacts data integrity in public databases and research findings.
- Workarounds and scripts are available to prevent or mitigate gene name conversion errors.