Related Experiment Video
Updated: Jul 12, 2025

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
Europe PMC annotated full-text corpus for gene/proteins, diseases and organisms
Xiao Yang1, Shyamasree Saha1,2, Aravind Venkatesan1
1Literature Services, EMBL-EBI, Wellcome Trust Genome Campus, Cambridge, UK.
Abstract:
Named entity recognition (NER) is a widely used text-mining and natural language processing (NLP) subtask. In recent years, deep learning methods have superseded traditional dictionary- and rule-based NER approaches. A high-quality dataset is essential to fully leverage recent deep learning advancements. While several gold-standard corpora for biomedical entities in abstracts exist, only a few are based on full-text research articles. The Europe PMC literature database routinely annotates Gene/Proteins, Diseases, and Organisms entities. To transition this pipeline from a dictionary-based to a machine learning-based approach, we have developed a human-annotated full-text corpus for these entities, comprising 300 full-text open-access research articles. Over 72,000 mentions of biomedical concepts have been identified within approximately 114,000 sentences. This article describes the corpus and details how to access and reuse this open community resource.
Related Concept Videos
Genomics
Genome Annotation and Assembly
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Protein Families
Animal Mitochondrial Genetics
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

