Related Experiment Video
Updated: Jan 22, 2026

10:16
Mining Spatial Transcriptomics Datasets using DeepSpaceDB
Published on: September 5, 2025
679
OryzaGP: rice gene and protein dataset for named-entity recognition
Pierre Larmande1,2, Huy Do2, Yue Wang3
1UMR DIADE, Institute of Research for Sustainable Development (IRD), F-34394 Montpellier, France.
Genomics & Informatics
|July 16, 2019
Summary
Researchers developed a new dataset for rice gene and protein recognition to improve text mining in plant molecular biology. This benchmark aids machine learning by enabling accurate analysis of rice scientific literature.
Area of Science:
- Plant Molecular Biology
- Bioinformatics
- Computational Biology
Background:
- Text mining is crucial for extracting biological entities like genes and proteins from scientific literature.
- Limited datasets and benchmarks exist for plant molecular biology, particularly for rice, hindering advanced analysis.
- Accurate named-entity recognition for rice is challenging due to the scarcity of specialized resources.
Purpose of the Study:
- To develop a new benchmark dataset for rice gene and protein name recognition.
- To facilitate the application of advanced machine learning methods for analyzing rice literature.
- To establish a shared task for evaluating text mining approaches in rice molecular biology.
Main Methods:
- Compiled a dataset of titles and abstracts from PubMed focusing on rice research.
- Utilized the 5th Biomedical Linked Annotation Hackathon for data sharing via PubAnnotation.
- Designed the dataset to serve as a benchmark for named-entity recognition tasks.
Main Results:
- Created a novel, publicly available dataset specifically for rice gene/protein entity extraction.
- The dataset addresses the lack of benchmarks for rice literature analysis.
- Facilitates comparative evaluation of different text mining and machine learning approaches.
Conclusions:
- The new rice dataset is essential for advancing text mining in plant molecular biology.
- It enables the development and evaluation of improved named-entity recognition systems for rice.
- Promotes open comparison and progress in analyzing rice scientific literature through shared tasks.
Related Concept Videos
Naming Enantiomers
25.7K
The naming of enantiomers employs the Cahn–Ingold–Prelog rules that involve assigning priorities to different substituent groups at a chiral center. Each enantiomer, being a distinct molecule, is assigned a unique name by the Cahn–Ingold–Prelog (CIP) rules, also called the R–S system. The prefix R- or S- attached to the chiral centers in an enantiomer is dependent on the spatial arrangement of the four substituents on the chiral center. The R–S system essentially comprises three...
25.7K
Proteins: From Genes to Degradation
14.2K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
Transcription is the synthesis of RNA...
14.2K
Proteins: From Genes to Degradation
4.3K
4.3K
Naming Skeletal Muscles
3.9K
The naming of the approximately 700 muscles in the human body is based on a set of criteria designed to provide descriptive information about each muscle, making it easier to identify and remember them.
The key factors used in naming muscles include:
The key factors used in naming muscles include:
3.9K
Common Names of Aldehydes and Ketones
4.9K
Some common aldehydes and ketones are popularly known by their common names used historically and predate the IUPAC nomenclature.
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the...
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the...
4.9K
Gene Conversion
10.6K
Other than maintaining genome stability via DNA repair, homologous recombination plays an important role in diversifying the genome. In fact, the recombination of sequences forms the molecular basis of genomic evolution. Random and non-random permutations of genomic sequences create a library of new amalgamated sequences. These newly formed genomes can determine the fitness and survival of cells. In bacteria, homologous and non-homologous types of recombination lead to the evolution of new...
10.6K

