Related Experiment Video
Updated: Sep 10, 2025

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
Published on: October 24, 2019
Unlocking the potential of PubMed Central supplementary data files
Julien Gobeill1,2, Déborah Caucheteur1,2, Alexandre Flament1,2
1SIB Text Mining Group, Swiss Institute of Bioinformatics, Geneva 1206, Switzerland.
Motivation:
Biocuration workflows often rely on comprehensive literature searches for specific biological entities. However, standard search engines such as MEDLINE and PubMed Central provide an incomplete picture of the scientific literature because they do not index the increasing amount of valuable information published in supplementary data files. Over two years, we addressed this gap by systematically extracting text from a large proportion (85%) of these files, resulting in 35 million searchable documents. To assess the information gain provided by supplementary data files beyond the manuscripts, we searched both for mentions of dozens of Global Core Biodata Resources (GCBRs), which are fundamental biological databases essential for the life sciences. We searched for mentions of GCBR names and accession numbers, which uniquely identify biological entities within these resources.
Results:
The recall gain from using the supplementary data files to search for articles mentioning resource names is 6%. In addition, 97% of all accession numbers identified were published in the supplementary data files, highlighting their increasing importance for highly specific topics or curation pipelines. We show that the number of accession numbers published in the supplementary data files is increasing year on year, but that 87% of these are published in Excel files. This format facilitates human readability and accessibility, but severely limits machine reusability and interoperability. We therefore discuss alternative and complementary approaches to the publication of research data.
Availability And Implementation:
All extracted data are accessible and searchable as a collection on the BiodiversityPMC platform (https://biodiversitypmc.sibils.org/).
More Related Videos
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
Globular and Fibrous Proteins
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Additional Subnuclear Structures
The nucleus contains many membrane-less subnuclear organelles or nuclear bodies, such as nucleoli, Cajal bodies, speckles,...
Western Blotting
The technique begins with separating proteins from the sample using sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), followed by protein transfer, immunoblotting, and finally, protein detection.
Analysis of Population Pharmacokinetic Data