Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Mar 28, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.3K

Automatic Entity Recognition and Typing from Massive Text Corpora: A Phrase and Network Mining Approach.

Xiang Ren1, Ahmed El-Kishky1, Chi Wang2

  • 1University of Illinois at Urbana-Champaign, Urbana, IL, USA.

KDD : Proceedings. International Conference on Knowledge Discovery & Data Mining
|December 26, 2015
PubMed
Summary

This tutorial introduces data-driven methods for recognizing typed entities in large text datasets. These techniques automatically identify and label entities, aiding knowledge discovery and management in diverse information sources.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Perfluorooctanesulfonic acid exposure leads to downregulation of 3-hydroxy-3-methylglutaryl-CoA synthase 2 expression and upregulation of markers associated with intestinal carcinogenesis in mouse intestinal tissues.

Chemosphere·2024
Same author

Overexpression of Fatty Acid Synthase Upregulates Glutamine-Fructose-6-Phosphate Transaminase 1 and O-Linked N-Acetylglucosamine Transferase to Increase O-GlcNAc Protein Glycosylation and Promote Colorectal Cancer Growth.

International journal of molecular sciences·2024
Same author

Factors associated with patients' healthcare-seeking behavior and related clinical outcomes under China's hierarchical healthcare delivery system.

Frontiers in public health·2024
Same author

Combined effect of time in target range and variability of systolic blood pressure on cardiovascular outcomes and mortality in patients with hypertension: A prospective cohort study.

Journal of clinical hypertension (Greenwich, Conn.)·2024
Same author

Effects of glycerol monolaurate on estradiol and follicle-stimulating hormones, offspring quality, and mRNA expression of reproductive-related genes of zebrafish (Danio rerio) females.

Fish physiology and biochemistry·2024
Same author

Nursing Staff Presenteeism Scale: Development and psychometric test.

PloS one·2024

Area of Science:

  • Computer Science
  • Information Science
  • Natural Language Processing

Background:

  • Modern society generates vast amounts of unstructured text data from diverse sources like news, social media, and scientific literature.
  • Extracting meaningful information and understanding relationships between entities within this data is crucial for knowledge discovery and management.

Purpose of the Study:

  • To introduce data-driven methods for recognizing typed entities in massive, domain-specific text corpora.
  • To demonstrate the scalability and effectiveness of these methods for identifying and labeling entities.

Main Methods:

  • Utilizing data-driven approaches to automatically identify token spans as entity mentions within documents.
  • Implementing methods for labeling identified entities with specific types (e.g., people, products, food).

Related Experiment Videos

Last Updated: Mar 28, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.3K
  • Applying these techniques to large-scale, domain-specific text collections.
  • Main Results:

    • Successful identification and typing of entities in real-world datasets, including news articles and tweets.
    • Demonstration of scalable entity recognition across different text domains.
    • Validation of the utility of typed entities for knowledge discovery and management.

    Conclusions:

    • Data-driven entity recognition is a powerful technique for unlocking value from unstructured text data.
    • The presented methods offer a scalable solution for identifying and categorizing entities in massive corpora.
    • Typed entities significantly enhance knowledge discovery and management processes in information-rich environments.