LARGE-SCALE MULTIPLE INFERENCE OF COLLECTIVE DEPENDENCE WITH APPLICATIONS TO PROTEIN FUNCTION
Robert Jernigan1, Kejue Jia1, Zhao Ren2
1Department of Biochemistry, Biophysics, and Molecular Biology, Program of Bioinformatics and Computational Biology, Iowa State University.
This study introduces collective dependence, a novel measure for analyzing complex relationships among multiple variables, particularly useful for understanding protein coevolution and identifying functional residue triplets.
Area of Science:
- Computational Biology
- Information Theory
- Statistical Inference
Background:
- Analyzing higher-order dependencies (k ≥ 3) among random variables is crucial but challenging.
- Protein coevolution, involving multivariate categorical features, requires advanced methods beyond pairwise analysis.
Purpose of the Study:
- To introduce and validate a novel information-theoretic measure, collective dependence, for quantifying higher-order variable interactions.
- To develop a robust statistical inference procedure for detecting significant k-collective dependence in large datasets.
- To apply the method to protein sequence data for uncovering novel coevolving residue insights.
Main Methods:
- Defined collective dependence as a symmetrized measure generalizing mutual information for k ≥ 3 variables.
- Developed a Classification-Assisted Large scaLe inference procedure (CALDECO) for detecting significant k-collective dependence with controlled false discovery rate.
- Validated the method through simulations and applied it to protein sequence alignment data.
Main Results:
- Collective dependence is easily estimated and facilitates dependence testing for k ≥ 3 variables.
- CALDEDECO successfully identified significant higher-order co-dependencies in simulated data.
- Novel functional triplets of amino acid residues were identified in elongation factor P and zinc knuckle protein families.
Conclusions:
- Collective dependence provides valuable insights into protein coevolution, surpassing traditional pairwise measures.
- The developed CALDECO procedure offers a powerful tool for large-scale inference of complex variable interactions.
- Identified coevolving residue triplets offer new avenues for investigating protein function and mechanisms.
More Related Videos
07:28JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
Protein Families
