Related Experiment Video
Updated: Apr 12, 2026

14:48
Endogenous Protein Tagging in Human Induced Pluripotent Stem Cells Using CRISPR/Cas9
Published on: August 25, 2018
29.0K
A hybrid named entity tagger for tagging human proteins/genes
Summary
This study introduces a novel hybrid approach for accurately identifying human gene and protein names in biomedical texts. The new method improves upon existing Named Entity Recognition (NER) systems, enhancing scientific literature analysis.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Natural Language Processing
Background:
- Accurate extraction of gene/protein names is crucial for biomedical literature analysis.
- Existing Named Entity Recognition (NER) tools often lack high performance for human gene/protein tagging.
- A precise tagger for human genes and proteins is essential for current text mining research.
Purpose of the Study:
- To develop a novel hybrid approach for improved human gene and protein name recognition in biomedical texts.
- To address limitations and common errors found in existing NER taggers.
- To enhance the accuracy of Named Entity Recognition (NER) for human gene/protein identification.
Main Methods:
- A hybrid approach combining machine learning (Conditional Random Fields), manually curated rules, and a novel abbreviation identification algorithm.
- Development of a domain-specific algorithm tailored for human gene and protein nomenclature.
- Utilizing the JNLPBA2004 corpus for experimental evaluation.
Main Results:
- The proposed hybrid approach achieved a precision of 80.47% and an F-score of 75.77% on the JNLPBA2004 corpus.
- The system demonstrated superior performance compared to most existing state-of-the-art NER systems for human gene/protein tagging.
- A recall of 71.60% was achieved, indicating areas for future methodological refinement.
Conclusions:
- The developed hybrid NER system offers a significant advancement in accurately identifying human genes and proteins.
- The approach effectively surmounts common errors in existing taggers, improving scientific literature analysis.
- Further research is needed to enhance recall and achieve even higher accuracy in gene/protein name extraction.
Related Concept Videos
Tagging and Fusion Proteins
8.9K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.9K
Hybridoma Technology
18.7K
Hybridoma technology is used for the large-scale production of monoclonal antibodies. Monoclonal antibodies bind to only a single antigenic determinant or epitope. Such antibodies are used in research, diagnostics, and disease therapy. The hybridoma technology established in 1975 by Georges Köhler and Cesar Milstein was awarded the Nobel Prize in Medicine in 1984 for revolutionizing research and therapy.
Hybridoma Selection
Commonly used fusion techniques — electroporation,...
Hybridoma Selection
Commonly used fusion techniques — electroporation,...
18.7K
RNA-seq
12.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.6K
Genome Annotation and Assembly
22.2K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
22.2K
Peptide Identification Using Tandem Mass Spectrometry
8.8K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.8K
Labeling DNA Probes
9.8K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
9.8K

