Related Experiment Video
Updated: May 25, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Automatic categorization of diverse experimental information in the bioscience literature
Ruihua Fang1, Gary Schindelman, Kimberly Van Auken
1Howard Hughes Medical Institute and Biology Division, California Institute of Technology, Pasadena, CA 91125, USA.
We developed an automated method using Support Vector Machines (SVM) to identify relevant scientific papers for biocuration, significantly reducing manual effort. This machine learning approach enhances the efficiency of biological knowledge database curation.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Curation
Background:
- Biocuration involves extracting experimental information from scientific literature into computable biological knowledge databases.
- A critical bottleneck in biocuration is the manual identification of relevant papers for specific data types, which is time-consuming.
- Existing methods require extensive manual review, leading to delays in knowledge base updates.
Purpose of the Study:
- To develop an automated method for identifying scientific papers relevant to specific biocuration data types.
- To reduce the time and labor involved in the manual screening of literature for biocuration.
Main Methods:
- Utilized the machine learning algorithm Support Vector Machine (SVM) for automated classification of scientific papers.
- Developed a procedure to leverage training papers from diverse literature sources for improved identification of low-occurrence data types.
- Implemented a system for automatic categorization of experimental data types.
Main Results:
- Successfully tested the SVM-based method on ten data types from WormBase, fifteen from FlyBase, and three from Mouse Genomics Informatics (MGI).
- The system is operational at WormBase, automatically associating new publications with ten data types.
- The method demonstrates applicability across diverse experimental data types and literature corpora.
Conclusions:
- The developed automated methods are effective for a variety of data types with training sets of several hundred to a few thousand documents.
- The system is fully automatic and can be integrated into various literature-based database workflows.
- This approach significantly contributes to automating the labor-intensive biocuration process, enhancing efficiency in biological knowledge management.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:11CorrelationCalculator and Filigree: Tools for Data-Driven Network Analysis of Metabolomics Data
Published on: November 10, 2023
Related Concept Videos
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Methods of Classification and Identification
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Evolutionary Relationships through Genome Comparisons