Related Experiment Video
Updated: Oct 17, 2025

14:58
High-Throughput Transcriptome Analysis for Investigating Host-Pathogen Interactions
Published on: March 5, 2022
4.5K
LitCovid-AGAC: cellular and molecular level annotation data set based on COVID-19
Sizhuo Ouyang1, Yuxing Wang1, Kaiyin Zhou1
1Hubei Key Lab of Agricultural Bioinformatics, College of Informatics, Huazhong Agricultural University, 430070 Wuhan, China.
Genomics & Informatics
|October 12, 2021
Summary
Researchers created a large, cross-annotated corpus of COVID-19 abstracts to aid knowledge discovery. This resource, LitCovid-AGAC, facilitates understanding the pathological mechanisms of coronavirus disease 2019.
Area of Science:
- Biomedical Natural Language Processing (BioNLP)
- Computational Biology
- Text Mining
Background:
- The dramatic increase in COVID-19 literature necessitates efficient text curation for knowledge discovery.
- BioNLP plays a crucial role in extracting information on COVID-19 mechanisms.
- Existing annotation systems like PubAnnotation facilitate data integration.
Purpose of the Study:
- To construct a comprehensive, cross-annotated corpus of COVID-19 abstracts for enhanced text mining.
- To integrate diverse annotation resources for a richer dataset.
- To facilitate the discovery of hidden knowledge regarding COVID-19's pathological mechanisms.
Main Methods:
- Merged three distinct annotation resources (PubTator, OGER, AGAC) into a unified dataset.
- Utilized the LitCovid dataset comprising 50,018 COVID-19 abstracts.
- Developed the LitCovid-AGAC corpus with 12 distinct labels (Mutation, Species, Gene, Disease, GO, CHEBI, Var, MPA, CPA, NegReg, PosReg, Reg).
Main Results:
- Successfully created the LitCovid-AGAC corpus, a cross-annotated resource for COVID-19 abstracts.
- The corpus integrates 12 types of annotations from multiple sources, providing rich information.
- The dataset is poised to support large-scale text mining for COVID-19 research.
Conclusions:
- The LitCovid-AGAC corpus represents a valuable resource for the BioNLP community.
- This integrated corpus enables deeper investigation into the pathological mechanisms of COVID-19.
- Facilitates advanced text mining and knowledge discovery from the growing body of COVID-19 research.
Related Concept Videos
Genome Annotation and Assembly
19.6K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.6K
Single Nucleotide Polymorphisms-SNPs
16.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
16.8K
Mutations in Microorganisms
182
Mutations are heritable changes in an organism’s genome involving alterations in the base sequence of DNA or RNA. These changes can influence cellular processes and phenotypic traits, potentially transforming the unaltered wild type into a mutant form. Such changes, termed forward mutations, are pivotal in shaping the genetic diversity of organisms.RNA viruses exhibit the highest mutation rates due to the absence of robust proofreading mechanisms during genome replication. In contrast,...
182
Calmodulin-dependent Signaling
5.4K
Calmodulin (CaM) is a calcium-binding protein in eukaryotes that controls various calcium-regulated cellular processes. It has four calcium-binding sites that bind calcium to form the calcium-calmodulin ( Ca2+-CaM) complex. GPCR stimulation increases the calcium levels in the cells that bind to CaM and induces a conformational change.
The Ca2+-CaM complex does not have enzymatic activity by itself. Instead, the complex binds downstream target proteins, including membrane proteins or enzymes,...
The Ca2+-CaM complex does not have enzymatic activity by itself. Instead, the complex binds downstream target proteins, including membrane proteins or enzymes,...
5.4K
RNA-seq
10.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.5K
Genomics
38.0K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
38.0K

