GAPSCORE: finding gene and protein names one word at a time

Jeffrey T Chang1, Hinrich Schütze, Russ B Altman

  • 1Department of Genetics, Stanford Medical Center, 300 Pasteur Drive, Lane L 301, Mail Code 5120, Stanford, CA 94305-5120, USA.

Summary

We developed GAPSCORE, a new method to identify gene and protein names in text. This statistical approach analyzes word appearance, morphology, and context for improved biological data extraction.

Related Concept Videos

Cis-regulatory Sequences02:02

Cis-regulatory Sequences

Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
Ribosome Profiling02:24

Ribosome Profiling

Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Gene Evolution - Fast or Slow?02:05

Gene Evolution - Fast or Slow?

The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
Cis-regulatory Sequences02:02

Cis-regulatory Sequences

Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
Prokaryotic Gene Structure and Organization01:28

Prokaryotic Gene Structure and Organization

Prokaryotic genomes exhibit a streamlined organization of coding and non-coding regions essential for gene expression and protein synthesis. While coding regions contain the genetic instructions for proteins or functional RNAs, non-coding regions regulate the precise transcription and translation of these genes.Coding Regions: Proteins and RNAsThe primary coding regions, known as structural genes, include sequences transcribed into messenger RNA (mRNA) and ultimately translated into...
Translation in Prokaryotes01:29

Translation in Prokaryotes

Prokaryote translation is a complex, highly coordinated process that converts genetic information from mRNA into functional proteins. It involves three stages: initiation, elongation, and termination, each facilitated by specific molecular components.Initiation of TranslationThe process begins with the assembly of the ribosomal subunits and initiation factors on the mRNA. In bacteria, the 30S ribosomal subunit recognizes the Shine-Dalgarno sequence in the mRNA, a conserved region upstream of...