Related Experiment Video
Updated: Jul 11, 2026

Using RNA-sequencing to Detect Novel Splice Variants Related to Drug Resistance in In Vitro Cancer Models
Published on: December 9, 2016
Disambiguating proteins, genes, and RNA in text: a machine learning approach
V Hatzivassiloglou1, P A Duboué, A Rzhetsky
1Department of Computer Science, Columbia University, 1214 Amsterdam Avenue, New York, NY 10027, USA. vh@cs.columbia.edu
This study introduces an automated system to classify biological terms like protein, gene, and mRNA in text. Achieving up to 85% accuracy, it uses machine learning and unsupervised training on molecular biology literature.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing in Biology
Background:
- Accurate identification and classification of biological entities (proteins, genes, mRNA) in scientific literature are crucial for knowledge extraction.
- Manual annotation is time-consuming and prone to errors, necessitating automated solutions.
- Existing methods may require extensive labeled data or struggle with contextual ambiguity.
Purpose of the Study:
- To develop and evaluate an automated system for assigning class labels (protein, gene, mRNA) to biological terms within free text.
- To explore the effectiveness of different machine learning algorithms and contextual feature definitions for disambiguation.
- To propose and implement a fully unsupervised method for generating training data.
Main Methods:
- Implementation of three distinct machine learning algorithms.
- Development of extended contextual features to improve term disambiguation.
- Utilization of a fully unsupervised approach for acquiring training examples.
- System training and evaluation on a large corpus of 9 million words from molecular biology journal articles.
Main Results:
- The automated system achieved high accuracy rates, reaching up to 85% in classifying biological terms.
- The proposed unsupervised method for obtaining training data proved effective.
- The examination of different machine learning algorithms and feature sets provided insights into optimal configurations.
Conclusions:
- The developed automated system offers an efficient and accurate solution for biological term classification in free text.
- The unsupervised training approach reduces the dependency on manually curated datasets.
- This system has the potential to significantly aid in large-scale biological knowledge discovery and data mining.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Proteins: From Genes to Degradation
Transcription is the synthesis of RNA molecules by RNA...
RNA Editing
Leaky Scanning
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...

