Related Experiment Video
Updated: Sep 29, 2025

09:40
Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
8.8K
Parallel sequence tagging for concept recognition
Lenz Furrer1,2, Joseph Cornelius3,2, Fabio Rinaldi4,5,6,7
1Department of Computational Linguistics, University of Zurich, Zurich, Switzerland.
BMC Bioinformatics
|March 25, 2022
Summary
We developed a novel parallel architecture for biomedical Named Entity Recognition (NER) and Normalisation (NEN), outperforming traditional serial pipelines. This approach effectively combines classifier strengths for improved accuracy in concept recognition.
Area of Science:
- Biomedical Natural Language Processing
- Computational Biology
- Bioinformatics
Background:
- Named Entity Recognition (NER) and Normalisation (NEN) are crucial for biomedical text mining.
- Traditional serial pipelines for NER and NEN are susceptible to error propagation.
- A parallel architecture offers a potential solution to overcome limitations of serial processing.
Purpose of the Study:
- To propose and evaluate a parallel architecture for joint Named Entity Recognition and Normalisation (NER/NEN).
- To model NER and NEN as sequence-labeling tasks operating directly on source text.
- To investigate various harmonisation strategies for merging predictions from parallel classifiers.
Main Methods:
- Developed a parallel architecture for NER and NEN, treating both as sequence-labeling tasks.
- Examined different strategies for harmonising predictions from NER and NEN classifiers.
- Evaluated the approach on the CRAFT corpus (Version 4), including 20 annotation sets.
Main Results:
- The proposed parallel system significantly outperformed the baseline pipeline system on all 20 annotation sets of the CRAFT corpus.
- Optimising harmonisation strategies for each annotation set further refined system performance.
- Demonstrated the effectiveness of combining strengths from independent NER and NEN classifiers.
Conclusions:
- The parallel architecture successfully combines the strengths of NER and NEN classifiers.
- Effective prediction harmonisation requires individual calibration for each annotation set.
- This approach balances existing knowledge with the identification of novel concepts.
Related Concept Videos
RNA-seq
10.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.5K
Next-generation Sequencing
93.2K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
93.2K
Peptide Identification Using Tandem Mass Spectrometry
7.0K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
7.0K
Tagging and Fusion Proteins
7.2K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
7.2K
Per-Unit Sequence Models
129
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
129
Modern Molecular Taxonomy
208
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
208

