Related Experiment Video
Updated: Feb 24, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Cell phenotypes in the biomedical literature: a systematic analysis and text mining corpus
Noam H Rotenberg1, Robert Leaman1, Rezarta Islamaj1
1Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD, USA.
Abstract:
The variety of cell phenotypes identified by single-cell technologies is rapidly expanding, yet this knowledge is dispersed across the scientific literature and incompletely represented in structured resources. We present the CellLink corpus, a manually annotated collection of over 22,000 mentions of human and mouse cell populations in recent journal articles, distinguishing specific cell phenotypes, heterogeneous cell populations, and vague cell populations, and linking to Cell Ontology (CL) terms as either exact or related matches, covering nearly half of the terms in the current CL. A systematic analysis reveals lineage-specific patterns in how authors utilize anatomical context, molecular signatures, functional roles, developmental stage, and other attributes in cell naming. We show that fine-tuning transformer-based models on CellLink yields strong performance for named entity recognition, while embedding-based approaches support zero-shot entity linking and distinguishing exact from related matches. We further demonstrate the utility of CellLink to expand and refine the chondrocyte branch of CL.

