Related Experiment Video
Updated: Mar 7, 2026

The Lambda Select cII Mutation Detection System
Published on: April 26, 2018
nala: text mining natural language mutation mentions
Juan Miguel Cejuela1,2, Aleksandar Bojchevski1,2, Carsten Uhlig1
1TUM, Department of Informatics, Bioinformatics & Computational Biology - i12, Garching, Munich, Germany.
Motivation:
The extraction of sequence variants from the literature remains an important task. Existing methods primarily target standard (ST) mutation mentions (e.g. 'E6V'), leaving relevant mentions natural language (NL) largely untapped (e.g. 'glutamic acid was substituted by valine at residue 6').
Results:
We introduced three new corpora suggesting named-entity recognition (NER) to be more challenging than anticipated: 28-77% of all articles contained mentions only available in NL. Our new method nala captured NL and ST by combining conditional random fields with word embedding features learned unsupervised from the entire PubMed. In our hands, nala substantially outperformed the state-of-the-art. For instance, we compared all unique mentions in new discoveries correctly detected by any of three methods (SETH, tmVar, or nala ). Neither SETH nor tmVar discovered anything missed by nala , while nala uniquely tagged 33% mentions. For NL mentions the corresponding value shot up to 100% nala -only.
Availability And Implementation:
Source code, API and corpora freely available at: http://tagtog.net/-corpora/IDP4+ .
Contact:
nala@rostlab.org.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
More Related Videos
Related Concept Videos
Mutations
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Mutations
Mutation, Gene Flow, and Genetic Drift
Viral Mutations
Mutations in Microorganisms
Point and Frameshift Mutations

