Related Experiment Video
Updated: Jun 11, 2026

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
A language independent acronym extraction from biomedical texts with hidden Markov models
IEEE Transactions on Bio-Medical Engineering
|July 14, 2010
Summary
This study introduces a Hidden Markov Model (HMM) approach for extracting acronyms and their meanings from text. This probabilistic method effectively handles ambiguity and noise, achieving high precision and recall in the biomedical domain.
Area of Science:
- Computational linguistics
- Bioinformatics
- Natural Language Processing
Background:
- Extracting acronyms and their definitions from unstructured text is challenging due to ambiguity and noise.
- Existing methods struggle with the complexities of acronyms in specialized domains like biomedicine.
Purpose of the Study:
- To propose and evaluate a novel probabilistic model for acronym and definition extraction.
- To address the challenges of ambiguity and noise in acronym identification using Hidden Markov Models (HMM).
Main Methods:
- Modeling the extraction process as a stochastic process using Hidden Markov Models (HMM).
- Utilizing characters of the acronym as states and text tokens as signals and observations.
- Calculating the most probable definition based on HMM output.
Main Results:
- Achieved high precision (93.50%) and recall (85.50%) in extracting acronyms and definitions.
- Demonstrated superior performance with the highest F1 score (89.40%) compared to other methods on a biomedical corpus.
- Successfully handled ambiguous and noisy acronym coinage.
Conclusions:
- The proposed HMM-based approach offers a robust solution for acronym extraction in challenging environments.
- Probabilistic modeling is effective in managing the inherent ambiguity and noise in acronym definition identification.
- This method shows significant potential for applications in biomedical text mining and information retrieval.
Related Concept Videos
Leaky Scanning
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Genetic Lingo
Overview
Multi-pass Transmembrane Proteins and β-barrels
In multi-pass transmembrane proteins, the polypeptide chain crosses the membrane more than once. The transmembrane polypeptide chain either forms an α-helix or β-strand structure. α-Helix containing multi-pass transmembrane proteins are ubiquitous, whereas β-strand containing ones are mainly found in gram-negative bacteria, mitochondria, and chloroplasts.
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as G-protein-linked receptors (GPCRs) and...
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as G-protein-linked receptors (GPCRs) and...

