Related Experiment Video
Updated: Oct 15, 2025

07:09
A Bioinformatics Pipeline for Investigating Molecular Evolution and Gene Expression using RNA-seq
Published on: May 28, 2021
9.9K
TAGOPSIN: collating taxa-specific gene and protein functional and structural information
Eshan Bundhoo1, Anisah W Ghoorah2, Yasmina Jaufeerally-Fakim1
1Department of Agricultural and Food Science, Faculty of Agriculture, University of Mauritius, Reduit, 80837, Mauritius.
BMC Bioinformatics
|October 24, 2021
Summary
TAGOPSIN is a new Java program that systematically retrieves and integrates data from seven major biological databases. This tool streamlines comparative genomics and protein structure studies by creating a unified data warehouse for easier analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Public biological databases offer vast information for gene, genome, and protein analysis.
- Current data retrieval across multiple databases is often unsystematic and repetitive.
- Need for efficient tools to integrate diverse biological data for research.
Purpose of the Study:
- To develop a command-line program for systematic data retrieval from multiple biological databases.
- To facilitate integrated analysis for comparative genomics and protein structure studies.
- To create a unified data resource for various biological applications.
Main Methods:
- Developed TAGOPSIN (TAxonomy, Gene, Ontology, Protein, Structure INtegrated), a Java-based command-line tool.
- Integrated data from seven popular public biological databases.
- Tested the program with diverse organisms including eukaryotes, prokaryotes, and viruses.
Main Results:
- TAGOPSIN enables rapid and systematic retrieval of organism-centered data.
- Successfully integrated data for Homo sapiens and human coronavirus, demonstrating broad applicability.
- Created a consolidated data warehouse for enhanced biological research.
Conclusions:
- TAGOPSIN provides efficient data integration, simplifying the manipulation of interconnected biological information.
- Facilitates interspecific comparative analyses of protein-coding genes and identification of taxa-specific protein structures.
- The program is available via GitHub under the GNU General Public License.
Related Concept Videos
Tagging and Fusion Proteins
7.3K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
7.3K
Gene Families
9.3K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.3K
Genome Annotation and Assembly
19.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.5K

