Related Experiment Video
Updated: Jun 19, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Construction of an annotated corpus to support biomedical information extraction
Paul Thompson1, Syed A Iqbal, John McNaught
1National Centre for Text Mining, Manchester Interdisciplinary Biocentre, University of Manchester, 131 Princess Street, Manchester, M1 7DN, UK. paul.thompson@manchester.ac.uk
Researchers developed a new annotation scheme and corpus (GREC) for gene regulation events in biomedical text. This resource aids in training information extraction tools for enhanced knowledge discovery in molecular biology.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Natural Language Processing
Background:
- Information Extraction (IE) is crucial for knowledge discovery in biomedical text mining.
- Understanding verb and nominalized verb behavior is key for IE.
- Annotated corpora are valuable for training IE components.
Purpose of the Study:
- To define a novel annotation scheme for sentence-bound gene regulation events.
- To create a comprehensive corpus (GREC) for training IE systems.
- To identify participants and assign semantic roles for gene regulation events.
Main Methods:
- Developed a new annotation scheme for gene regulation events, including verbs and nominalized verbs.
- Created the Gene Regulation Event Corpus (GREC) with 240 annotated MEDLINE abstracts.
- Annotated participants with 13 tailored semantic roles and linked to the Gene Regulation Ontology.
Main Results:
- The GREC corpus includes detailed annotations for gene regulation and expression events.
- Annotation achieved inter-annotator agreement rates between 66% and 90%.
- The annotation scheme uniquely captures a wide range of event arguments in the biomedical field.
Conclusions:
- GREC is a unique biomedical resource annotating core relationships and contextual details.
- GREC supports the development of bio-specific tools and resources, including BioLexicon.
- The corpus is suitable for training IE components like semantic role labellers and is freely available.
More Related Videos
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025