Related Experiment Video
Updated: May 14, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
The Metadata Coverage Index (MCI): A standardized metric for quantifying database metadata richness
Konstantinos Liolios1, Lynn Schriml, Lynette Hirschman
1Microbial Genomics and Metagenomic Super Program, Department of Energy Joint Genome Institute, Walnut Creek, CA, USA.
Public data repositories often have inconsistent metadata quality, making individual record assessment impractical. We introduce the Metadata Coverage Index (MCI), a simple metric to objectively score and filter data records, driving improvements in metadata richness and data analysis.
Area of Science:
- Bioinformatics
- Data Science
- Scientific Data Management
Background:
- Public data repositories contain variable metadata, necessitating impractical individual quality assessments by users.
- Objective measures are needed to evaluate and improve the richness of data descriptions for better data usability.
Purpose of the Study:
- To introduce the Metadata Coverage Index (MCI) as a quantitative metric for assessing data record description completeness.
- To demonstrate the utility of MCI for filtering, ranking, and improving metadata quality in public repositories.
Main Methods:
- The Metadata Coverage Index (MCI) is defined as the percentage of available fields filled within a data record.
- MCI scores were calculated and analyzed using metadata from the Genomes Online Database (GOLD), including records adhering to the Minimum Information about a Genome Sequence (MIGS) standard.
Main Results:
- The MCI provides an objective proxy for metadata quality, enabling practical filtering and analysis of large datasets.
- Demonstrated MCI's utility in assessing metadata completeness and identifying areas for improvement within the GOLD database.
Conclusions:
- The MCI offers a valuable, objective framework for quantifying metadata coverage and driving improvements in data annotation practices.
- This metric can inform standards bodies, repository providers, and curators, ultimately enhancing the usability and value of public data.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Methods to Assess Microbial Communities
Midrange
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to outliers and...
Central Tendency: Analysis
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
Chemical Ionization (CI) Mass Spectrometry
Magnetic Resonance Imaging
