Related Experiment Videos
Literature consistency of bioinformatics sequence databases is effective for assessing record quality
Summary
Bioinformatics sequence databases contain millions of records, but many have data quality issues. This study introduces quality indicators to detect literature inconsistencies, finding one in four records may be faulty, enabling automated error detection.
Area of Science:
- Bioinformatics
- Genomic Data Analysis
- Scientific Literature Mining
Background:
- Genomic sequence databases like GenBank and UniProt house millions of records.
- These large-scale datasets suffer from data quality issues such as errors, redundancies, and inconsistencies with published literature.
- Curators face challenges in identifying anomalous and suspicious records within these vast databases.
Purpose of the Study:
- To investigate and analyze the data quality of genomic sequence databases, with a focus on detecting literature inconsistencies.
- To develop and validate a set of quality indicators for assessing record-literature consistency.
- To explore the potential for automated detection of faulty records based on literature inconsistencies.
Main Methods:
- Proposed a set of 24 quality indicators based on querying the published literature for each record.
- Analyzed the mutual relationship between proposed quality indicators and overall record quality.
- Utilized Principal Component Analysis (PCA) for dimensionality reduction and visualization of record-literature consistency vectors.
- Manually analyzed records identified as potentially inconsistent based on their proximity to known erroneous records.
Main Results:
- A mutual dependency was observed between the proposed quality indicators and the quality of the records.
- Records with literature inconsistencies were visualized in a similar region after dimensionality reduction using PCA.
- Manual analysis revealed that approximately one in four records is inconsistent with the published literature.
- A high density of inconsistent records suggests potential for automated detection methods.
Conclusions:
- Literature inconsistency is a significant and meaningful strategy for identifying suspicious records in bioinformatics databases.
- The developed quality indicators and analysis framework provide a basis for improving data quality in genomic sequence databases.
- The findings highlight the need for enhanced data curation and validation processes in large-scale bioinformatics resources.