Related Experiment Video
Updated: May 12, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
tmVar: a text mining approach for extracting sequence variants in biomedical literature
Chih-Hsuan Wei1, Bethany R Harris, Hung-Yu Kao
1National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), 8600 Rockville Pike, Bethesda, MD 20894, USA.
tmVar is a novel text-mining tool that accurately extracts diverse mutation types from biomedical literature. This bioinformatics approach improves upon existing methods for analyzing sequence variations in complex diseases.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Automated text-mining of mutation data is crucial for understanding sequence variations in complex diseases.
- Existing methods often focus on limited mutation types (e.g., protein point mutations) and require extensive manual rule development.
- There is a need for accurate, automatic approaches to extract a wider range of mutation information.
Purpose of the Study:
- To develop and evaluate tmVar, a text-mining approach for extracting diverse sequence variants from biomedical literature.
- To cover important mutation types not previously addressed by existing text-mining tools.
- To improve the accuracy and scope of mutation extraction compared to state-of-the-art methods.
Main Methods:
- Utilized a conditional random field (CRF) model for text mining.
- Developed a novel CRF label model and feature set tailored for mutation extraction.
- Applied the approach to protein, DNA, and RNA sequence variants adhering to Human Genome Variation Society nomenclature.
Main Results:
- tmVar achieved high performance, outperforming a state-of-the-art method on both a custom corpus and a gold-standard dataset.
- F-measure scores for tmVar were 91.4% (custom corpus) and 93.9% (gold standard), compared to 78.1% and 89.4% respectively for the comparison method.
- The approach successfully extracts a wide range of mutation types, including those not covered by previous studies.
Conclusions:
- tmVar is a high-performance tool for extracting diverse mutation information from biomedical literature.
- The developed CRF model and feature set significantly enhance mutation extraction accuracy.
- This tool facilitates the analysis of sequence variations and supports the creation of mutation databases.
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Sanger Sequencing
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...

