Related Experiment Video
Updated: Oct 27, 2025

06:41
In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
13.9K
Computational predictions for protein sequences of COVID-19 virus via machine learning algorithms
Heba M Afify1, Muhammad S Zanaty2
1Systems and Biomedical Engineering Department, Higher Institute of Engineering in El-Shorouk City, Cairo, Egypt. hebaaffify@yahoo.com.
Medical & Biological Engineering & Computing
|July 22, 2021
Summary
This study classifies COVID-19 protein sequences by country using machine learning. The model achieved 100% accuracy, offering a predictive tool for global viral protein analysis.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning in virology
Background:
- The COVID-19 pandemic necessitates understanding viral protein variations across human populations.
- Global spread of SARS-CoV-2 impacts different countries uniquely.
Purpose of the Study:
- To classify human protein sequences of COVID-19 based on country of origin.
- To develop a machine learning model for predicting geographic protein sequence patterns.
Main Methods:
- Utilized conjoint triad (CT) method for data preprocessing, converting amino acids to numerical groups.
- Implemented two data labeling strategies: country code assignment and binary element representation.
- Employed machine learning algorithms, including linear Support Vector Machine (SVM), for sequence classification.
Main Results:
- Achieved 100% accuracy and sensitivity, with 90% specificity using binary labeling and linear SVM.
- Demonstrated higher classification accuracy for countries with more available protein sequence data, such as the USA.
- Identified data imbalance as a significant challenge, with US data comprising 76% of the total sequences.
Conclusions:
- The developed machine learning model effectively classifies COVID-19 protein sequences by country.
- The model shows potential as a predictive tool for understanding geographic variations in viral proteins.
- Addressing data imbalance is crucial for improving the model's generalizability across all countries.
Related Concept Videos
Conjugated Proteins
21.8K
Simple proteins and protein complexes contain only amino acids. In contrast, many other proteins, called conjugated proteins, covalently bond with non-protein moieties.
Nucleoproteins are protein complexes that contain nucleic acids, categorized as deoxyribonucleoproteins (DNPs) or ribonucleoproteins (RNPs) respectively. The nucleosome is a typical example of a DNP where nuclear DNA is associated with histone proteins. The major antigen for the Covid-19 virus SARS-CoV is an RNP that is critical...
Nucleoproteins are protein complexes that contain nucleic acids, categorized as deoxyribonucleoproteins (DNPs) or ribonucleoproteins (RNPs) respectively. The nucleosome is a typical example of a DNP where nuclear DNA is associated with histone proteins. The major antigen for the Covid-19 virus SARS-CoV is an RNP that is critical...
21.8K
Protein Families
16.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.2K
Single Nucleotide Polymorphisms-SNPs
17.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.0K
Conserved Binding Sites
4.7K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.7K
Protein-protein Interfaces
14.1K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.1K

