Related Experiment Video
Updated: Jul 6, 2026

10:41
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved (Non-model) Organisms
Published on: May 9, 2017
Natural language processing in aid of FlyBase curators
Nikiforos Karamanis1, Ruth Seal, Ian Lewin
1Computer Laboratory, University of Cambridge, William Gates Building, Cambridge, CB3 0FD, UK. nikiforos.karamanis@cl.cam.ac.uk
BMC Bioinformatics
|April 16, 2008
Summary
Natural Language Processing (NLP) tools like PaperBrowser enhance biomedical database curation. This NLP-powered interface improves navigational efficiency by 58% and utility by 74% for curators.
Area of Science:
- Biomedical Informatics
- Computational Linguistics
Background:
- Growing interest in applying Natural Language Processing (NLP) to biomedical text.
- Uncertainty regarding NLP's effectiveness in facilitating tasks like database curation.
Purpose of the Study:
- To develop and evaluate PaperBrowser, an NLP-powered interface designed to improve article navigation for FlyBase curators.
- To assess the impact of a user-centered design approach on the effectiveness of biomedical text curation tools.
Main Methods:
- User-centered design informed by observing curators at work.
- User-based study evaluating PaperBrowser's navigational functionalities using a text highlighting task.
- Assessment based on Human-Computer Interaction (HCI) criteria and navigational efficiency metrics.
Main Results:
- PaperBrowser significantly improves navigational efficiency, reducing interactions between highlighting events by approximately 58%.
- The interface enhances navigational utility for curators by over 74%, regardless of individual highlighting techniques.
- Demonstrates the practical application of NLP in streamlining complex curation workflows.
Conclusions:
- State-of-the-art NLP tasks, including Named Entity Recognition and Anaphora Resolution, can be effectively integrated with PaperBrowser's navigation features.
- PaperBrowser successfully supports biomedical database curation by improving efficiency and utility.
- Highlights the potential of NLP-driven tools in advancing biomedical research support systems.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Genetic Screens
Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
FISH - Fluorescent In-situ Hybridization
Fluorescence in situ hybridization, or FISH, was developed in the early 1980s and has quickly become one of the most widely used techniques in cytogenetics. Labeled probes are used to bind complementary DNA or RNA sequences on a chromosome or in a region within a cell. Earlier, the probes could only be obtained by cloning or reverse transcription of a DNA template. Currently, the probe oligonucleotides can be synthesized synthetically. Additionally, with the advancement of optical techniques,...
Cis-regulatory Sequences
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
Nonsense-mediated mRNA Decay
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...

