Related Experiment Video
Updated: Feb 16, 2026

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
Published on: October 24, 2019
Scholarly Information Extraction Is Going to Make a Quantum Leap with PubMed Central (PMC)
1Jena University Language & Information Engineering (JULIE) Lab, Friedrich-Schiller-Universität Jena, Jena 07743, Germany.
Abstract:
With the increasing availability of complete full texts (journal articles), rather than their surrogates (titles, abstracts), as resources for text analytics, entirely new opportunities arise for information extraction and text mining from scholarly publications. Yet, we gathered evidence that a range of problems are encountered for full-text processing when biomedical text analytics simply reuse existing NLP pipelines which were developed on the basis of abstracts (rather than full texts). We conducted experiments with four different relation extraction engines all of which were top performers in previous BioNLP Event Extraction Challenges. We found that abstract-trained engines loose up to 6.6% F-score points when run on full-text data. Hence, the reuse of existing abstract-based NLP software in a full-text scenario is considered harmful because of heavy performance losses. Given the current lack of annotated full-text resources to train on, our study quantifies the price paid for this short cut.
Related Concept Videos
Quantum Numbers
The Quantum-Mechanical Model of an Atom
The Central Dogma
The Central Dogma
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
Measures of Central Tendency
Trait Centrality

