CB-Search: a method for searching for similar protein motifs with biased compositions
1Department of Computer Networks and Systems, Silesian University of Technology, Akademicka 2A, 44-100, Gliwice, Poland. patryk.jarnot@polsl.pl.
Background:
Methods for analyzing protein sequence similarity focus mainly on identifying the homologies of proteins with standard amino acid compositions. For motifs with biased compositions, they have already been shown to be suboptimal; thus, we lack dedicated tools to support their analyses. However, motifs with compositional biases also play key roles in protein functions. These domains can be found in transmembrane proteins, bind to RNA through RGG boxes, and may form prions. Nevertheless, many domains remain unknown, as for a long time they were considered nonfunctional and most of the methods mask them to improve homology searches. Therefore, we need better solutions to infer their functions more efficiently.
Results:
In this research, we developed a new method incorporating three alignment strategies, an algorithm for identifying motifs with similar compositions, 2-mer based filtering, and a new metric for evaluating alignments. These solutions focus mainly on comparing the physicochemical properties of protein sequences rather than their evolutionary relationships. To validate our approach, we compared BLAST with our method in three variants that use local, global-local, and global alignment with the algorithm for identifying compositionally similarities. We used these methods to search for similar transmembrane domains and RGG boxes. We observed that our solutions significantly increased the number of true positives. The greatest increase occurred after we applied our similarity score measure. Compositionally biased motifs frequently consist of two adjacent functionally important motifs; therefore, we also searched for similarities to the K-DE motif of DNA-directed RNA polymerase subunit delta. We found that compared with the other alignment strategies, global-local and global alignment with identifying similar regions included all the submotifs of the query sequence more often.
Conclusion:
Our method introduces novel strategies that enhance the search for compositionally biased motifs, thereby improving annotation retrieval via sequence matching.
Related Concept Videos
Protein Families
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...


