Related Experiment Video
Updated: Feb 5, 2026

06:33
Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
7.2K
Integrating Language Model and Reading Control Gate in BLSTM-CRF for Biomedical Named Entity Recognition
IEEE/ACM Transactions on Computational Biology and Bioinformatics
|September 6, 2018
Summary
This study introduces a novel Long Short Term Memory (LSTM) Networks model for biomedical named entity recognition (Bio-NER). The enhanced LS-BLSTM-CRF model achieves superior performance, outperforming existing state-of-the-art systems.
Area of Science:
- Computational Biology
- Bioinformatics
- Natural Language Processing
Background:
- Biomedical Named Entity Recognition (Bio-NER) is crucial for biomedical text mining.
- Current neural network methods for Bio-NER overlook sentence-level semantic and general linguistic features.
- Complex hand-designed features are often avoided in mainstream NER approaches.
Purpose of the Study:
- To propose a novel Long Short Term Memory (LSTM) Networks model for Bio-NER.
- To integrate sentence-level semantic information and general linguistic features into Bio-NER models.
- To improve the accuracy and performance of Bio-NER systems.
Main Methods:
- Developed a Long Short Term Memory (LSTM) Networks model integrating a language model and a sentence-level reading control gate (LS-BLSTM-CRF).
- Incorporated a sentence-level reading control gate (SC) to capture implicit sentence meaning.
- Utilized character-level embeddings to handle out-of-vocabulary words.
Main Results:
- The proposed LS-BLSTM-CRF model achieved an F-score of 89.94% on the BioCreative II GM corpus.
- The model outperformed all existing state-of-the-art Bio-NER systems.
- Achieved a 1.33% improvement over the best-performing neural networks.
Conclusions:
- The novel LS-BLSTM-CRF model effectively integrates sentence-level semantics and language model features for Bio-NER.
- The proposed approach significantly enhances Bio-NER performance compared to current methods.
- Character-level embeddings are beneficial for handling unseen words in biomedical text mining.
Related Concept Videos
Naming Enantiomers
26.1K
The naming of enantiomers employs the Cahn–Ingold–Prelog rules that involve assigning priorities to different substituent groups at a chiral center. Each enantiomer, being a distinct molecule, is assigned a unique name by the Cahn–Ingold–Prelog (CIP) rules, also called the R–S system. The prefix R- or S- attached to the chiral centers in an enantiomer is dependent on the spatial arrangement of the four substituents on the chiral center. The R–S system essentially comprises three...
26.1K
Language
918
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
918
Naming Skeletal Muscles
4.1K
The naming of the approximately 700 muscles in the human body is based on a set of criteria designed to provide descriptive information about each muscle, making it easier to identify and remember them.
The key factors used in naming muscles include:
The key factors used in naming muscles include:
4.1K
Uncertainty in Measurement: Reading Instruments
52.9K
Counting is the type of measurement that is free from uncertainty, provided the number of objects being counted does not change during the process. Such measurements result in exact numbers. By counting the eggs in a carton, for instance, one can determine exactly how many eggs are there in the carton. Similarly, the numbers of defined quantities are also exact. For example, 1 foot is exactly 12 inches, 1 inch is exactly 2.54 centimeters, and 1 gram is exactly 0.001 kilograms. Quantities...
52.9K
Common Names of Aldehydes and Ketones
5.0K
Some common aldehydes and ketones are popularly known by their common names used historically and predate the IUPAC nomenclature.
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the...
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the...
5.0K
Components of Language
821
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
821

