OryzaGP: rice gene and protein dataset for named-entity recognition

Pierre Larmande1,2, Huy Do2, Yue Wang3

  • 1UMR DIADE, Institute of Research for Sustainable Development (IRD), F-34394 Montpellier, France.

Summary

Researchers developed a new dataset for rice gene and protein recognition to improve text mining in plant molecular biology. This benchmark aids machine learning by enabling accurate analysis of rice scientific literature.

Related Concept Videos

Naming Enantiomers02:21

Naming Enantiomers

The naming of enantiomers employs the Cahn–Ingold–Prelog rules that involve assigning priorities to different substituent groups at a chiral center. Each enantiomer, being a distinct molecule, is assigned a unique name by the Cahn–Ingold–Prelog (CIP) rules, also called the R–S system. The prefix R- or S- attached to the chiral centers in an enantiomer is dependent on the spatial arrangement of the four substituents on the chiral center. The R–S system essentially comprises three...
25.7K
Proteins: From Genes to Degradation02:11

Proteins: From Genes to Degradation

Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick.  Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
14.2K
Proteins: From Genes to Degradation02:11

Proteins: From Genes to Degradation

4.3K
Naming Skeletal Muscles01:19

Naming Skeletal Muscles

The naming of the approximately 700 muscles in the human body is based on a set of criteria designed to provide descriptive information about each muscle, making it easier to identify and remember them.
The key factors used in naming muscles include:
3.9K
Common Names of Aldehydes and Ketones01:11

Common Names of Aldehydes and Ketones

Some common aldehydes and ketones are popularly known by their common names used historically and predate the IUPAC nomenclature.   
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the...
4.9K
Gene Conversion02:08

Gene Conversion

Other than maintaining genome stability via DNA repair, homologous recombination plays an important role in diversifying the genome. In fact, the recombination of sequences forms the molecular basis of genomic evolution. Random and non-random permutations of genomic sequences create a library of new amalgamated sequences. These newly formed genomes can determine the fitness and survival of cells. In bacteria, homologous and non-homologous types of recombination lead to the evolution of new...
10.6K