Related Experiment Video
Updated: Jul 4, 2025

13:56
A Comparative Approach to Characterize the Landscape of Host-Pathogen Protein-Protein Interactions
Published on: July 18, 2013
11.2K
Multi-modal features-based human-herpesvirus protein-protein interaction prediction by using LightGBM
Xiaodi Yang1, Stefan Wuchty2,3,4,5, Zeyin Liang1
1Department of Hematology, Peking University First Hospital, Beijing, China.
Briefings in Bioinformatics
|January 27, 2024
Summary
This study introduces a novel multi-modal embedding fusion method using LightGBM to accurately predict human-herpesvirus protein-protein interactions (PPIs). The approach enhances understanding of viral infections, particularly in cancer patients.
Area of Science:
- Virology
- Bioinformatics
- Computational Biology
Background:
- Identifying human-herpesvirus protein-protein interactions (PPIs) is crucial for understanding viral infection mechanisms, especially in cancer patients.
- Existing natural language processing (NLP) methods for predicting human-herpesvirus PPIs are limited, particularly in multi-modal feature fusion.
- Herpesvirus infections are common in malignant tumor patients, necessitating advanced methods for studying viral-host interactions.
Purpose of the Study:
- To develop and evaluate a multi-modal embedding feature fusion method for predicting human-herpesvirus PPIs.
- To improve the accuracy and comprehensiveness of PPI prediction by integrating sequence, network, and functional data.
- To assess the model's generalizability across different herpesvirus subtypes using transfer learning.
Main Methods:
- A LightGBM model was developed, integrating multi-modal features from human and herpesviral proteins.
- Document and graph embedding approaches were used to represent sequence, network, and functional characteristics.
- Models were trained and validated on rigorous and non-rigorous benchmarking datasets, comparing performance against individual modal features and other machine learning methods.
Main Results:
- The multi-modal embedding fusion method significantly outperformed models using individual modal features.
- The developed LightGBM model demonstrated superior performance compared to traditional feature encoding methods and state-of-the-art deep learning approaches.
- Transfer learning showed the model could reliably predict human-cytomegalovirus PPIs even without cytomegalovirus-specific data, indicating broad applicability.
Conclusions:
- Multi-modal feature fusion using LightGBM is a powerful strategy for predicting human-herpesvirus PPIs.
- The method effectively captures complex protein interaction features across diverse herpesvirus subtypes.
- This approach offers a valuable tool for studying viral infections and host-pathogen interactions in clinical settings.

