Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genetic Lingo01:11

Genetic Lingo

Overview
Genomics02:02

Genomics

Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Gene Flow02:39

Gene Flow

Gene flow is the transfer of genes among populations, resulting from either the dispersal of gametes or from the migration of individuals.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Assembly of a high-quality reference genome for the rat tapeworm <i>Hymenolepis diminuta</i>.

bioRxiv : the preprint server for biology·2026
Same author

Recognition and linking of discontinuous named entities in healthcare: a comparative performance analysis.

Frontiers in digital health·2026
Same author

Health inequalities in outpatient neurological conditions across a large UK urban population: a retrospective observational study using automated coding.

BMJ neurology open·2026
Same author

Intrinsic coordination of dynamic molecular signatures shape the human prefrontal cortex.

bioRxiv : the preprint server for biology·2026
Same author

Development and validation of prediction models for predicting social care strengths and vulnerability in older people: Cohort study using routine data in Adult Social Care.

PloS one·2026
Same author

An integrative single-nucleus multiomic atlas of the human left ventricle identifies gene regulatory network dynamics across cardiac development, aging, and disease.

Genome biology·2026

Related Experiment Video

Updated: May 30, 2026

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
09:37

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information

Published on: August 15, 2019

The GNAT library for local and remote gene mention normalization.

Jörg Hakenberg1, Martin Gerner, Maximilian Haeussler

  • 1Pharma Research and Early Development, Hoffmann-La Roche Inc., Nutley, NJ 07110, USA. jorg.hakenberg@roche.com

Bioinformatics (Oxford, England)
|August 5, 2011
PubMed
Summary

Gnat is a Java library for identifying and normalizing gene and protein names in biomedical text. This tool aids text mining and data analysis by improving named entity recognition and normalization for biological research.

More Related Videos

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

Related Experiment Videos

Last Updated: May 30, 2026

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
09:37

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information

Published on: August 15, 2019

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Natural Language Processing

Background:

  • Named entity recognition and normalization are crucial for biomedical text mining.
  • Limited availability of public tools hinders progress in this field.
  • The Gnat Java library addresses this gap by providing a robust solution.

Purpose of the Study:

  • To introduce the Gnat Java library for biomedical text processing.
  • To enable efficient identification and normalization of gene and protein mentions.
  • To offer a flexible tool for integration into existing text-mining pipelines.

Main Methods:

  • Development of a Java library for text retrieval, named entity recognition, and normalization.
  • Implementation of capabilities for gene and protein mention identification.
  • Evaluation on the BioCreative III test dataset.

Main Results:

  • The Gnat library effectively performs named entity recognition and normalization for gene and protein mentions.
  • Achieved a Tap-20 score of 0.1987 on the BioCreative III test data.
  • The library is available as source code and can be used as a stand-alone application or integrated component.

Conclusions:

  • Gnat provides a valuable resource for biomedical text mining and data analysis.
  • The library enhances the ability to process and extract information about genes and proteins from scientific literature.
  • Public availability of Gnat promotes further research and development in the field.