Related Experiment Video
Updated: Jul 7, 2026

22:10
Multi-target Parallel Processing Approach for Gene-to-structure Determination of the Influenza Polymerase PB2 Subunit
Published on: June 28, 2013
Rule-based knowledge aggregation for large-scale protein sequence analysis of influenza A viruses
Olivo Miotto1, Tin Wee Tan, Vladimir Brusic
1Institute of Systems Science, National University of Singapore, 25 Heng Mui Keng Terrace, Singapore. olivo@nus.edu.sg
BMC Bioinformatics
|March 20, 2008
Summary
Automated metadata extraction from biological databases improves data quality for large-scale analyses. This approach successfully annotates influenza A sequences, reducing manual curation needs and enhancing data reliability.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Science
Background:
- The rapid expansion of biological data necessitates advanced methods for analyzing large sequence alignments.
- Ensuring metadata quality is crucial for reliable results in comparative biological studies.
- Semantic heterogeneity and inconsistencies in public databases complicate metadata aggregation and cleaning.
Purpose of the Study:
- To investigate quality issues in major public biological databases.
- To quantify the effectiveness of an automated metadata extraction approach.
- To annotate influenza A sequences with key properties like protein name, virus subtype, host, and isolation details.
Main Methods:
- Applied an automated metadata extraction approach combining structural and semantic rules.
- Processed over 90,000 influenza A records from NCBI public databases.
- Utilized user-defined structural rules for data aggregation and reconciliation.
Main Results:
- Extracted and reconciled metadata for over 40,000 influenza A protein sequences with >88.8% recovery and >96% accuracy.
- Identified significant quality differences between databases, with GenBank yielding more reliable values than GenPept.
- Reconstructed isolate relationships, identified 7,640 isolates, and corrected over 3,000 metadata inconsistencies.
Conclusions:
- Automated knowledge aggregation with embedded intelligence is essential for large-scale biological data analysis.
- User-controlled rule-based approaches can automate curation tasks, reducing manual corrections to ~5%.
- Semantic technologies offer significant potential for improving knowledge aggregation in bioinformatics.
More Related Videos
Related Concept Videos
Leaky Scanning
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Influenza
Influenza is an acute, highly communicable viral disease that affects the respiratory tract and is responsible for seasonal epidemics worldwide. Influenza A is the most prevalent type associated with widespread outbreaks and is subtyped based on two surface glycoproteins: hemagglutinin (H) and neuraminidase (N), as in H1N1. These glycoproteins are essential for viral infectivity, transmission, and immune recognition. Transmission occurs primarily through respiratory droplets and contaminated...
Viral Mutations
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material for adaptive...
Viruses with RNA Genomes
RNA viruses are categorized into positive-strand, negative-strand, or double-stranded groups based on their genomic structure and replication mechanisms. This classification dictates how they exploit host cellular machinery for protein synthesis and replication. Some RNA viruses also utilize reverse transcription as part of their life cycle, further diversifying their replication strategies.Positive-Strand RNA VirusesPositive-strand RNA viruses have genomes that function directly as messenger...
Inhibitors Of Virion Release
Viral replication and dissemination rely on efficient mechanisms for host cell entry, genome replication, assembly, and release. Influenza viruses, such as types A and B, are negative-sense single-stranded RNA viruses with a segmented genome, that depend on two critical surface glycoproteins to carry out these processes: hemagglutinin (HA) and neuraminidase (NA). HA initiates infection by binding to sialic acid residues on the surface of host epithelial cells, facilitating receptor-mediated...

