Related Experiment Videos
Bio-Ontology and text: bridging the modeling gap.
Carol Friedman1, Tara Borlawsky, Lyudmila Shagina
1Department of Biomedical Informatics, Columbia University, New York, NY 10032, USA. Friedman@dbmi.columbia.edu
Bioinformatics (Oxford, England)
|July 28, 2006
Summary
We developed PGschema, a novel translational schema to integrate biological text data with genetic databases. This schema effectively represents phenotypic and genetic information, enabling advanced systems biology research.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Natural Language Processing (NLP) is increasingly used to extract biological discoveries from text.
- Textual data structures differ significantly from ontology-based databases.
- Integrating unstructured text with genetic databases is crucial for systems biology.
Purpose of the Study:
- To propose and evaluate a translational schema for representing phenotypic and genetic information from natural language.
- To connect biological information across different scales.
- To map textual information to existing biological ontologies for better data integration.
Main Methods:
- Developed PGschema, a novel representational schema.
- Designed the schema to translate phenotypic, genetic, and related information from text.
- Included concepts from established ontologies, modifiers, and relationships.
Main Results:
- PGschema successfully represented 90% of selected entities (95% CI: 86-93%).
- The schema can be automatically expressed in XML format using NLP techniques.
- This is the first evaluation of a translational schema for NLP containing declarative knowledge on genes and phenotypes.
Conclusions:
- PGschema facilitates the computational reuse and integration of biological information from text.
- The schema is essential for managing heterogeneous phenotypic information and advancing systems biology.
- It enables the acquisition and reuse of diverse knowledge from large volumes of text.