Related Experiment Video
Updated: Jul 11, 2026

08:43
Metagenomic Analysis of Silage
Published on: January 13, 2017
ARC: automated resource classifier for agglomerative functional classification of prokaryotic proteins using
Muthiah Gnanamani1, Naveen Kumar, Srinivasan Ramachandran
1G N Ramachandran Knowledge Centre for Genome Informatics, Institute of Genomics and Integrative Biology, Mall Road, Delhi 110 007, India.
Journal of Biosciences
|October 5, 2007
Summary
Automated Resource Classifier (ARC) is a flexible, open-source software for protein functional classification. It aids comparative genomics by accurately categorizing proteins into functional classes with high success rates.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Functional classification of proteins is crucial for comparative genomics.
- Interpreting analytical data requires flexible, automated algorithms.
- A general, automated software solution is needed to aid this process.
Purpose of the Study:
- To develop ARC (Automated Resource Classifier), an open-source software for flexible protein functional classification.
- To provide a general, automated tool for integrative interpretation of analytical data in genomics.
Main Methods:
- ARC utilizes a keyword-matching, agglomerative classification scheme.
- The keyword library was built using data from Bacillus subtilis, Escherichia coli K12, Archaeoglobus fulgidus, Gene Ontology, and Gene Symbols.
- The software classifies proteins into 7 basic and 2 ancillary functional classes.
Main Results:
- ARC achieved 94.04% success on 675,663 annotated proteins from 348 prokaryotes.
- The software demonstrates flexibility and meets user requirements.
- Examples illustrate its application in understanding mycobacterial physiology and protein costs.
Conclusions:
- ARC is a highly successful and flexible tool for automated protein functional classification.
- The software significantly aids comparative genomics and data interpretation.
- ARC is publicly available for researchers.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Aggregates Classification
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Tagging and Fusion Proteins
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
Microbial Classification System
Classification is the process of organizing organisms into hierarchically inclusive groups based on their phenotypic similarities or evolutionary relationships. A species comprises one or more strains, and closely related species are grouped into genera. Genera are further classified into families, families into orders, orders into classes, and so forth, up to the domain level, which is the broadest taxonomic rank derived from a combination of phenotypic and genotypic data.The nomenclature of...

