Related Experiment Video
Updated: Jul 9, 2025

A Web Tool for Generating High Quality Machine-readable Biological Pathways
Published on: February 8, 2017
A machine learning-enabled open biodata resource inventory from the scientific literature.
Heidi J Imker1,2, Kenneth E Schackart1,3, Ana-Maria Istrate4
1Global Biodata Coalition, Strasbourg, France.
A new automated method using machine learning and natural language processing identified 3112 biodata resources from scientific literature, aiding the Global Biodata Coalition in understanding research infrastructure support.
Area of Science:
- Bioinformatics
- Data Science
- Life Sciences
Background:
- Modern biological research relies heavily on biodata resources for data archiving, curation, and analysis.
- Sustained funding for this global infrastructure of biodata resources presents a significant challenge.
- The Global Biodata Coalition (GBC) aims to develop sustainable funding strategies for these essential resources.
Purpose of the Study:
- To develop an automated method for creating a comprehensive global inventory of biodata resources.
- To overcome limitations of existing registries that require manual curation and self-registration.
- To enable periodic updates to the inventory with minimal human intervention.
Main Methods:
- Utilized machine learning-enabled natural language processing (NLP) on Europe PMC publication data.
- Fine-tuned Bidirectional Encoder Representations from Transformers (BERT) models for resource identification and name prediction.
- Incorporated manual review for low-confidence predictions and duplicate resolution, supplemented by article metadata for funder and geolocation data.
Main Results:
- Generated an inventory of 3112 unique biodata resources from articles published between 2011 and 2021.
- The NLP approach successfully identified biodata resources from publication titles and abstracts.
- Developed automated pipelines for reproducible inventory generation, with all code and data released under permissive licenses.
Conclusions:
- The automated NLP-driven approach provides an efficient and scalable method for cataloging biodata resources.
- This inventory aids the Global Biodata Coalition in assessing the scope and support of the global biodata infrastructure.
- The open release of code and data promotes reuse and supports the sustainability of vital biodata resources for biological research.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:37Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019