Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Recombinant DNA01:09

Recombinant DNA

97.0K
Overview
97.0K
pre-mRNA Processing02:01

pre-mRNA Processing

53.9K
In eukaryotic cells, transcripts made by RNA polymerase are modified and processed before exiting the nucleus. Unprocessed RNA is called precursor mRNA or pre-mRNA to distinguish it from mature mRNA.
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a “cap” to the 5’ end of the growing transcript. In this process, a 5’ phosphate is replaced by modified guanosine that has a methyl group attached to it (7-Methyl...
53.9K
Bioequivalence: Overview01:16

Bioequivalence: Overview

1.3K
Pharmaceutical equivalents, by definition, are drug products with the same active ingredient in the same quantities, encapsulated in identical dosage forms, and intended for the same administration routes. These pharmaceutical equivalents are deemed bioequivalent if the bioavailability of the active entity in the drug preparations is similar. Moreover, pharmaceutical equivalents demonstrating bioequivalence are also regarded as therapeutically equivalent. This means that when used as directed,...
1.3K
MicroRNAs01:22

MicroRNAs

22.0K
MicroRNA (miRNA) are short, regulatory RNA transcribed from introns—non-coding regions of a gene—or intergenic regions—stretches of DNA present between genes. Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After...
22.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.8K
2.8K
CRISPR01:59

CRISPR

53.2K
Genome editing technologies allow scientists to modify an organism’s DNA via the addition, removal, or rearrangement of genetic material at specific genomic locations. These types of techniques could potentially be used to cure genetic disorders such as hemophilia and sickle cell anemia. One popular and widely used DNA-editing research tool that could lead to safe and effective cures for genetic disorders is the CRISPR-Cas9 system. CRISPR-Cas9 stands for Clustered Regularly Interspaced...
53.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Microbial named entity recognition and normalisation for AI-assisted literature review and meta-analysis.

Bioinformatics (Oxford, England)·2026
Same author

A maturity model framework for federated networks of trusted research environments.

Frontiers in digital health·2026
Same author

Erratum: Intestinal microbiome changes in response to amino acid and micronutrient supplementation: secondary analysis of the AMAZE trial - CORRIGENDUM.

Gut microbiome (Cambridge, England)·2026
Same author

AI-Assisted Pneumonia Detection, Localisation and Report Generation from Chest X-rays.

medRxiv : the preprint server for health sciences·2026
Same author

End-to-end integrative segmentation and radiomics prognostic models for risk stratification of high-grade serous ovarian cancer: a retrospective multicohort study.

The Lancet. Digital health·2026
Same author

From description to implementation: key takeaways from the 3rd African Microbiome Symposium.

mSphere·2025

Related Experiment Video

Updated: Oct 1, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

642

Auto-CORPus: A Natural Language Processing Tool for Standardizing and Reusing Biomedical Literature.

Tim Beck1,2, Tom Shorter1, Yan Hu3,4

  • 1Department of Genetics and Genome Biology, University of Leicester, Leicester, United Kingdom.

Frontiers in Digital Health
|March 4, 2022
PubMed
Summary

Auto-CORPus standardizes biomedical literature by converting HTML to BioC format and extracting tables and abbreviations. This novel Natural Language Processing tool supports text analytics and data sharing for researchers.

Keywords:
biomedical literaturehealth datanatural language processingsemanticstext mining

More Related Videos

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.0K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

8.9K

Related Experiment Videos

Last Updated: Oct 1, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

642
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.0K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

8.9K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Natural Language Processing

Background:

  • Standardizing biomedical literature is crucial for machine learning and Natural Language Processing (NLP) analysis.
  • Limited access to biomedical data in BioC format and a lack of tools to convert HTML hinder analysis.
  • Existing BioC specifications lack structures for publication table data.

Purpose of the Study:

  • Introduce Auto-CORPus, a novel NLP tool for standardizing and converting research publications.
  • Enable machine-interpretable data outputs to support biomedical text analytics.
  • Facilitate data sharing and analysis of diverse biomedical literature.

Main Methods:

  • Developed Auto-CORPus, an automated pipeline for converting publication HTML and table image files.
  • Utilized Information Artifact Ontology for annotating publication sections within BioC output.
  • Created a JSON format for representing publication table data and extracted abbreviations.

Main Results:

  • Auto-CORPus converts HTML to BioC format, standardizing publication sections with ontology annotations.
  • Publication tables from inline and linked sources are converted to a machine-interpretable JSON format.
  • Extracted abbreviations from publications, providing a JSON output of definitions for text mining.

Conclusions:

  • Auto-CORPus provides a novel solution for standardizing biomedical literature into machine-readable formats.
  • The tool enhances biomedical text analytics by enabling BioC conversion, table data processing, and abbreviation extraction.
  • Auto-CORPus supports data sharing and analysis, addressing limitations in current bioinformatics tools.